All Multimodal tools (56)
- UI-TARS-desktopByteDance's open-source multimodal AI agent stack for controlling computers and browsers, with Agent TARS CLI, remote operators, and the UI-TARS-1.5 model.
- MiniMaxShanghai lab behind Hailuo, M2.5, and the 2026 M3/H3 coding models. TalkAPI M3 is $0.24/$0.98 per 1M tokens.
- NotebookLMGoogle's source-grounded research notebook, now presented as Gemini Notebook at notebook.google. Analyzes your files and turns them into audio, video, and more.
- ToAPIsOpenAI-compatible AI API gateway: one base URL and key for 60+ models from GPT, Claude, Gemini, DeepSeek, Sora, and VEO, with routing, failover, and unified.
- DoubaoByteDance's consumer AI assistant. Chat, writing, coding, and Seedance video in one app. Official homepage currently offers Seedance 2.0 for free after login.
- TRAEByteDance's AI IDE. Recurring seats: Lite $3, Pro $10, Pro+ $30, Ultra $100. Usage is billed in dollars, not extra-fast request packs.
- Veo 3Google DeepMind's Veo video model, now presented as Veo 3.1: cinematic video with native audio, plus 1080p and 4K output on the official model page.
- YouMindYouMind is an AI creation studio that turns saved articles, videos, and notes into writing, slides, images, webpages, and video, with a paid Skills marketplace.
- Agnes-3.0-FlashAgnes-3.0-Flash Preview is an Apache-2.0 33B multimodal open-weight checkpoint with 262k context, not the proprietary API model.
- cogvlm-base-490-hfPrevious CogVLM snapshot. Not a 2026 VLM default.
- DeepSeek-V4.1-FlashDeepSeek-V4.1-Flash is a MIT 552B multimodal MoE with 8B/16B activation, 1M context, and about 1/4 the KV cache of V4-Flash.
- DeepSeek-V4-Flash-Vision-ExpDeepSeek-V4-Flash-Vision-Exp is the first V4 multimodal experiment: MIT vision weights on Flash, 305B params, screenshots plus text agents.
- deepseek-vl-7b-basePrevious DeepSeek-VL 7B Base snapshot. Not a 2026 VL default.
- FLUX 3Black Forest Labs multimodal model for video, audio, upcoming images, and action-prediction, with up to 20-second clips, native audio, and draft mode.
- Gemma 4 26B A4BGoogle DeepMind's open-weight 25.2B MoE model with 3.8B active parameters, 256K context, multimodal input, and a commercially permissive Apache 2.0 license.
- GLM-5.3-FlashZhipu's first natively multimodal GLM (confirmed as the 'Ox Alpha' stealth model): 320B MoE with 18B active, 1M context, MIT open weights at $0.15/M input.
- Google: Gemini 2.0 FlashPrevious Gemini 2.0 Flash snapshot. Live Flash is Gemini 3.7 Flash.
- Google: Gemini 3.1 Flash-LiteGoogle Gemini 3.1 Flash-Lite: leftover cheap Flash-Lite sibling on the live models page.
- Google: Gemini 3.1 Flash LiveGoogle Gemini 3.1 Flash Live: leftover live/audio sibling on the Gemini models page. Not the 2026 Flash default.
- Google: Gemini 3.1 ProGoogle Gemini 3.1 Pro: leftover Pro sibling on the live models page. Not a 2026 Pro default.
- Google: Gemini 3.5 Flash-LiteGoogle Gemini 3.5 Flash-Lite: leftover cheap Flash-Lite row on the live Gemini models page.
- Google: Gemini 3.5 FlashGoogle's agent-first frontier model: long-horizon agentic tasks, coding, and a 1M-token context at Flash-tier speed and cost.
- Google: Gemini 3.6 FlashGoogle Gemini 3.6 Flash: previous-generation Flash after 3.5, before live 3.7 Flash.
- Google: Gemini 3.7 FlashGoogle's most capable Flash model for agentic workflows and multimodal reasoning, with 1M-token context, 65k output tokens, and Computer Use preview.
- Google: Gemini 3 FlashGoogle's latest frontier model delivering breakthrough intelligence at unprecedented speed and cost efficiency.
- Google: Gemini 3 ProPrevious Gemini 3 Pro Preview snapshot. Not a 2026 vision crown.
- Google: Gemini Omni FlashGoogle Gemini Omni Flash: leftover Omni Flash sibling on the live Gemini models page. Not the 2026 Flash default.
- GPT Image 2OpenAI's latest image generation model for fast, high-quality create and edit jobs, with flexible sizes up to 4K and token-based API pricing.
- Grok 4.6SpaceXAI's frontier 1.5T reasoning model with 500K-token context, agentic coding strength, and real-time X data at roughly half the price of rivals.
- GrokxAI / SpaceXAI Grok family. Current docs flagship is grok-4.6: 500k context, $2/$6 per 1M tokens, plus Imagine image/video and Voice APIs.
- Jina Embeddings v4Previous Jina Embeddings v4 snapshot.
- Kimi K3Moonshot AI's open-weight 2.8T multimodal agentic model with 1M-token context, the world's first open 3T-class model rivaling closed frontier models.
- Ling-3.0-flash-VLLing-3.0-flash-VL is inclusionAI's MIT multimodal MoE: 124B total, 5.5B active, with image and video.
- llava-v1.6-34b-hfPrevious LLaVA-NeXT / LLaVA 1.6 snapshot.
- Meta Llama 3.2 Vision2024 Llama 3.2 vision snapshot. Not the 2026 multimodal default.
- MiniMax H3MiniMax open multimodal video model (July 31, 2026): text/image/video/audio in, 768P or 2K out, 4-15s clips, native stereo, pay-as-you-go API.
- MiniMax M3MiniMax open-weight frontier model (June 1, 2026): 1M-token MSA context, native image/video input, desktop computer use, and discounted API pricing.
- Mistral OCR 4.1Mistral OCR 4.1: live latest OCR with paragraph boxes, structural labels, and confidence. Not leftover OCR 4.0.
- Mistral Pixtral 12B2024 Mistral 12B vision snapshot. Not the first-and-only multimodal.
- Mistral Shieldstral 1.0Mistral Shieldstral 1.0: live compact multimodal moderation model. Apache 2.0.
- Nex-N2.5-miniNex-N2.5-mini is Nex-AGI's Apache-2.0 multimodal agent model for computer use, browsing, and coding.
- Nex-N2.5-ProNex-N2.5-Pro is Nex-AGI's Apache-2.0 multimodal agent checkpoint for computer use, browsing, and coding.
- OpenAI: GPT-4o-miniPrevious GPT-4o mini snapshot. Not the 2026 latest.
- OpenAI: GPT-4oPrevious GPT-4o snapshot. Not the 2026 latest.
- Qwen-Drive-1.0-4BQwen-Drive-1.0-4B is Alibaba's Apache-2.0 driving VLM that unifies 3D perception, VQA, and motion planning.
- Qwen-VLPrevious Alibaba Qwen-VL snapshot. Not a 2026 LVLM default.
- Qwen2-VL 72B InstructPrevious Qwen2-VL 72B Instruct snapshot. Not a 2024 VL crown.
- Qwen3.8-2.4T-A95BAlibaba Qwen3.8 2.4T MoE with 95B active, 1M context, multimodal input. Coder $0.14/$0.80; 397B $0.29/$1.20.
- Qwen3.8-27BAlibaba's open-weight 27B dense companion to Qwen3.8-Max: Apache 2.0 license, multimodal image-text input, built for local deployment and small-batch inference.
- Qwen3.8-Flash-NextAlibaba's open-weight architecture preview of Qwen4: 125B multimodal MoE with 6B active plus a 51B N-gram table, 262K native context, at $0.16/M input.
- Qwen3.8-MaxAlibaba's 2.4T-parameter open-weight flagship: 95B active MoE, text/image/video input, 1M context, at $2/$6 per million tokens.
- Qwen3-VL-EmbeddingPrevious Qwen3-VL embedding snapshot.
- Qwen3-VL-RerankerPrevious Qwen3-VL rerank snapshot.
- Seedance 2.5ByteDance Seed's audio-video model for 30-second storytelling, precise reference control, extend-twice generation, and pro tools like green-screen editing.
- WeMM-EmbeddingTencent WeChat Vision's universal multimodal embedding family (2B/4B/9B) built on Qwen3.5. Text, image, video, and visual-document inputs map to one L2-normalized embedding space.
- YuE2-3BYuE2-3B is M·A·P's CC BY-NC music model that plans an editable score, then renders vocals and accompaniment locally.
All Tags
A2A ProtocolAgent FrameworkAgent OrchestrationAgent RuntimeAgent WalletAI AgentAI AssistantAI ContentAI PaymentsAI SearchAlibabaAnthropicAPIAutomationAWSBAAIBlack Forest LabsBrowser AutomationByteDanceChatbotClaudeClaude CodeCLICloudCloudflareCode AssistantCode QualityCode ReviewCodexCoding AgentCohereCommunicationComputer UseCrypto & Web3CursorData ScienceData VisualizationDatabaseDebuggingDeep ResearchDeepgramDeepSeekDeploymentDesignDeveloper ToolsDocument ProcessingDocumentationEducationElevenLabsEmbeddingEnterpriseEvaluationFine-TuningFLUXFrontier ModelGeminiGemmaGitGoGoogleGPTGPUGrokHealthHugging FaceIBMImage GenerationIntegrationJina AIKimiKnowledge BaseLangChainLangGraphLinuxLlamaLocal AILong ContextmacOSMarkdownMCPMetaMicrosoftMiniMaxMistralMixedbread AIMixture of ExpertsModel RoutingMoonshot AIMulti-AgentMultilingualMultimodal
Items tagged with Multimodal (56)
Month Visit20000000DR: 90AS: 95
Month Visit20000000DR: 90AS: 95
Month Visit15700000DR: 72AS: 58
Month Visit15000000DR: 95AS: 98
Month Visit12000000DR: 88AS: 90
Month Visit12000000DR: 88AS: 90
Month Visit4800000DR: 95AS: 80
Month Visit2500000DR: 70AS: 60
Month Visit800000DR: 90AS: 88
Month Visit800000DR: 90AS: 88
Month Visit800000DR: 78AS: 82
Month Visit500000DR: 65AS: 70
Month Visit150000DR: 60AS: 55
Month Visit1500DR: 95AS: 92
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit800DR: 88AS: 85
Month Visit215DR: 92AS: 87
Month Visit245DR: 91AS: 85