All Mixture of Experts tools (19)
- DeepSeek V3Previous DeepSeek V3 snapshot. Not a GPT-4o / Claude 3.5 census.
- DeepSeek-V4.1-FlashDeepSeek-V4.1-Flash is a MIT 552B multimodal MoE with 8B/16B activation, 1M context, and about 1/4 the KV cache of V4-Flash.
- DeepSeek-V4-Flash-Vision-ExpDeepSeek-V4-Flash-Vision-Exp is the first V4 multimodal experiment: MIT vision weights on Flash, 305B params, screenshots plus text agents.
- DeepSeek V4 FlashDeepSeek small agent model: 284B/13B MoE, 1M context, MIT weights, $0.14/$0.28 per 1M tokens.
- DeepSeek V4DeepSeek V4 represents the next generation of DeepSeek's flagship AI models, building upon the success of V3 with enhanced capabilities in reasoning, multimodal processing, and agent-based interactions.
- Gemma 4 26B A4BGoogle DeepMind's open-weight 25.2B MoE model with 3.8B active parameters, 256K context, multimodal input, and a commercially permissive Apache 2.0 license.
- GLM-5.3-FlashZhipu's first natively multimodal GLM (confirmed as the 'Ox Alpha' stealth model): 320B MoE with 18B active, 1M context, MIT open weights at $0.15/M input.
- Hy4 previewTencent's next-generation open-weight MoE flagship: 770B total params, 49B activated per token, 1M-token context. Apache 2.0, Gated DSA attention, and blind-eval scores edging out GLM 5.3 and Kimi K3.
- KAT-Coder V2.5Kwaipilot's open-weight agentic coding model: 35B MoE, 3B active per token, Qwen3.6 base, Apache 2.0, 262K context, top PinchBench tool-use score.
- Laguna S 2.1Poolside's open-weight 118B MoE coding model with 8B active parameters, a 1M-token context window, native interleaved reasoning, and an OpenMDW-1.1 license.
- Ling-3.0-flash-VLLing-3.0-flash-VL is inclusionAI's MIT multimodal MoE: 124B total, 5.5B active, with image and video.
- LongCat 2.0Meituan's open-weight MoE model: 1.6T total / ~48B active params, 1M context, MIT. Strong on coding and agentic tasks (2026-08-25).
- MiniMax M2.1Open-source 230B parameter MoE model optimized for multi-language coding, agentic workflows, and real-world development tasks with 74% SWE-bench performance.
- NVIDIA Nemotron 3.5 Lightning 30B A3BNVIDIA's efficient open-weight 30B MoE hybrid model with 3B active parameters, 1M-token context, and single-GPU deployment for local reasoning and coding.
- Qwen3.8-Flash-NextAlibaba's open-weight architecture preview of Qwen4: 125B multimodal MoE with 6B active plus a 51B N-gram table, 262K native context, at $0.16/M input.
- Qwen3.8-MaxAlibaba's 2.4T-parameter open-weight flagship: 95B active MoE, text/image/video input, 1M context, at $2/$6 per million tokens.
- Tiel-Coder 35B A3BAgentic coding MoE. Not a lab release.
- WeLM-617BWeChat AI's 617B MoE model with 23B active parameters, demonstrating the first sequence-length scaling at frontier scale via Hidden Decoding, powering WeChat.
- WizardLM-2 8x22BPrevious WizardLM-2 8x22B snapshot.
All Tags
A2A ProtocolAgent FrameworkAgent OrchestrationAgent RuntimeAgent WalletAI AgentAI AssistantAI ContentAI PaymentsAI SearchAlibabaAnthropicAPIAutomationAWSBAAIBlack Forest LabsBrowser AutomationByteDanceChatbotClaudeClaude CodeCLICloudCloudflareCode AssistantCode QualityCode ReviewCodexCoding AgentCohereCommunicationComputer UseCrypto & Web3CursorData ScienceData VisualizationDatabaseDebuggingDeep ResearchDeepgramDeepSeekDeploymentDesignDeveloper ToolsDocument ProcessingDocumentationEducationElevenLabsEmbeddingEnterpriseEvaluationFine-TuningFLUXFrontier ModelGeminiGemmaGitGoGoogleGPTGPUGrokHealthHugging FaceIBMImage GenerationIntegrationJina AIKimiKnowledge BaseLangChainLangGraphLinuxLlamaLocal AILong ContextmacOSMarkdownMCPMetaMicrosoftMiniMaxMistralMixedbread AIMixture of Experts
Items tagged with Mixture of Experts (19)
Month Visit45000000DR: 95AS: 88
Month Visit500000DR: 80AS: 75
Month Visit500000DR: 68AS: 72
Month Visit345DR: 100AS: 100
Month Visit100DR: 100AS: 100
Month Visit135DR: 95AS: 88
Month Visit135DR: 95AS: 88
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80
Month Visit1000DR: 80AS: 80