AREX-2
AREX-2 is a 27-billion-parameter agent model from Beijing Academy of Artificial Intelligence, released on Hugging Face on 2026-09-29. It is described as a long-horizon agent model that learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise. The model is built on a Qwen3.8-compatible dense multimodal base, uses Apache-2.0, and supports a 262,144-token context. Training focuses on machine-learning and algorithmic programming tasks with verifiable feedback, and the learned behavior also transfers to deep research.
Key Features
- Long-horizon self-improvement: Uses extra test-time rounds to refine solutions rather than treating iteration as cost with no benefit.
- Feedback-driven reflection: Reads scores, logs, errors, and timings to decide what to change next.
- Cross-domain transfer: Coding and machine-learning training also improves deep-research behavior without new search trajectories.
- 27B dense multimodal architecture: Text and image input support with a 262,144-token context.
- Open license and weights: Apache-2.0 weights are available directly on Hugging Face.
Use Cases
Who Should Use This Tool?
- Researchers: A useful open checkpoints for studying long-horizon agent loops and self-improvement.
- ML engineers: A 27B model that can be run on a single workstation with modest serving requirements compared with much larger agents.
- Deep research integrators: Teams that want a text-and-image agent with long context and iterative reasoning.
Problems It Solves
- Long task horizons: Many models degrade as the number of rounds grows; AREX-2 is trained to keep productive refinement going.
- Verifiable coding and ML work: It can treat logs, scores, and errors as feedback signals.
- Cost of large closed agents: The 27B open checkpoint can be self-hosted and avoids API token costs for certain workloads.
Pricing
| Plan | Price | Features |
|---|---|---|
| Open weights | $0 software | Apache-2.0 checkpoint on Hugging Face; you pay for compute and hosting. |
Advantages & Unique Selling Points
Compared to Competitors: At 27B it is far smaller than frontier API agents, but the project positions it specifically for long-horizon iteration and deep research rather than general chatbot use.
What Makes It Stand Out: The public model card reports strong results on Frontier-CS, MLE-Lite, BrowseComp, GAIA, and DeepSearchQA in its own evaluation protocols. Those numbers are self-reported and should be treated as vendor claims, not independent audits.
Getting Started
- Install a recent Transformers release with Qwen3.8 support.
- Load
BAAI/AREX-2withAutoModelForMultimodalLMandAutoProcessorin bfloat16. - Send a prompt that asks for a proposed solution and an explanation of how it would be improved over several rounds.
- Use a GPU or multi-GPU setup because 27B parameters in bfloat16 need meaningful VRAM.
Integration
- Hugging Face for model weights.
- Python Transformers for inference.
- Qwen3.8-compatible tokenizer and processor.
Frequently Asked Questions
Is AREX-2 open source?
The weights are Apache-2.0 on Hugging Face and the project links a public GitHub page.
How big is the context window?
262,144 tokens according to the model card.
Does it accept images?
Yes. The card lists it as a dense Qwen3.8-compatible multimodal model.
Are the benchmark numbers independent?
No. The table in the model card is the project's own evaluation and should be viewed as a first-party claim.
Alternatives
If AREX-2 is not the right fit, consider these alternatives:
- Qwen3.8-Max: A much larger open-weight flagship with multimodal input and broad availability.
- DeepSeek-V4-Flash: A more widely benchmarked lightweight model family if you need proven local coding performance.
- MiniMax H3: A video model with broad availability when you need a different production-scale checkpoint.
Tips & Best Practices
- Test AREX-2 on tasks where you can give it quantitative feedback; its advantage is iterative refinement.
- Budget VRAM carefully: 27B in bfloat16 needs substantial GPU memory.
- Treat the model-card evaluations as directional evidence and run your own benchmark before choosing it over a frontier API.
Conclusion
AREX-2 is a compact open-weight model aimed at one specific behavior: improving solutions over many rounds. It is a research-minded checkpoint that fits teams wanting long-context agent reasoning without paying for a closed frontier service. Confirm the current inference requirements and benchmark protocols before committing a production workflow to it.
Comments
No comments yet. Be the first to comment!
Related Tools
Related Insights

Anthropic Subagent: The Multi-Agent Architecture Revolution
Deep dive into Anthropic multi-agent architecture design. Learn how Subagents break through context window limitations, achieve 90% performance improvements, and real-world applications in Claude Code.
Stop Cramming AI Assistants into Chat Boxes: Clawdbot Picked the Wrong Battlefield
Clawdbot is convenient, but putting it inside Slack or Discord was the wrong design choice from day one. Chat tools are not for operating tasks, and AI isn't for chatting.

Grok Bot and Hermes Bot: one person finally gets a think tank and a secretariat
Grok Bot now ships with Cursor Pro+. Hermes Bot runs on a VPS. They are not smarter chat boxes. The think tank advises, the secretariat executes, and you still make the call.