AREX-2 logo

AREX-2

Visit

BAAI's 27B open-weight agent model for long-horizon self-improvement: reflect, measure, and revise across coding, ML engineering, and deep research with a 262K context.

Share:
View alternatives

AREX-2

AREX-2 is a 27-billion-parameter agent model from Beijing Academy of Artificial Intelligence, released on Hugging Face on 2026-09-29. It is described as a long-horizon agent model that learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise. The model is built on a Qwen3.8-compatible dense multimodal base, uses Apache-2.0, and supports a 262,144-token context. Training focuses on machine-learning and algorithmic programming tasks with verifiable feedback, and the learned behavior also transfers to deep research.

Key Features

  • Long-horizon self-improvement: Uses extra test-time rounds to refine solutions rather than treating iteration as cost with no benefit.
  • Feedback-driven reflection: Reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain transfer: Coding and machine-learning training also improves deep-research behavior without new search trajectories.
  • 27B dense multimodal architecture: Text and image input support with a 262,144-token context.
  • Open license and weights: Apache-2.0 weights are available directly on Hugging Face.

Use Cases

Who Should Use This Tool?

  • Researchers: A useful open checkpoints for studying long-horizon agent loops and self-improvement.
  • ML engineers: A 27B model that can be run on a single workstation with modest serving requirements compared with much larger agents.
  • Deep research integrators: Teams that want a text-and-image agent with long context and iterative reasoning.

Problems It Solves

  1. Long task horizons: Many models degrade as the number of rounds grows; AREX-2 is trained to keep productive refinement going.
  2. Verifiable coding and ML work: It can treat logs, scores, and errors as feedback signals.
  3. Cost of large closed agents: The 27B open checkpoint can be self-hosted and avoids API token costs for certain workloads.

Pricing

Plan Price Features
Open weights $0 software Apache-2.0 checkpoint on Hugging Face; you pay for compute and hosting.

Advantages & Unique Selling Points

Compared to Competitors: At 27B it is far smaller than frontier API agents, but the project positions it specifically for long-horizon iteration and deep research rather than general chatbot use.

What Makes It Stand Out: The public model card reports strong results on Frontier-CS, MLE-Lite, BrowseComp, GAIA, and DeepSearchQA in its own evaluation protocols. Those numbers are self-reported and should be treated as vendor claims, not independent audits.

Getting Started

  1. Install a recent Transformers release with Qwen3.8 support.
  2. Load BAAI/AREX-2 with AutoModelForMultimodalLM and AutoProcessor in bfloat16.
  3. Send a prompt that asks for a proposed solution and an explanation of how it would be improved over several rounds.
  4. Use a GPU or multi-GPU setup because 27B parameters in bfloat16 need meaningful VRAM.

Integration

  • Hugging Face for model weights.
  • Python Transformers for inference.
  • Qwen3.8-compatible tokenizer and processor.

Frequently Asked Questions

Is AREX-2 open source?

The weights are Apache-2.0 on Hugging Face and the project links a public GitHub page.

How big is the context window?

262,144 tokens according to the model card.

Does it accept images?

Yes. The card lists it as a dense Qwen3.8-compatible multimodal model.

Are the benchmark numbers independent?

No. The table in the model card is the project's own evaluation and should be viewed as a first-party claim.

Alternatives

If AREX-2 is not the right fit, consider these alternatives:

  • Qwen3.8-Max: A much larger open-weight flagship with multimodal input and broad availability.
  • DeepSeek-V4-Flash: A more widely benchmarked lightweight model family if you need proven local coding performance.
  • MiniMax H3: A video model with broad availability when you need a different production-scale checkpoint.

Tips & Best Practices

  1. Test AREX-2 on tasks where you can give it quantitative feedback; its advantage is iterative refinement.
  2. Budget VRAM carefully: 27B in bfloat16 needs substantial GPU memory.
  3. Treat the model-card evaluations as directional evidence and run your own benchmark before choosing it over a frontier API.

Conclusion

AREX-2 is a compact open-weight model aimed at one specific behavior: improving solutions over many rounds. It is a research-minded checkpoint that fits teams wanting long-context agent reasoning without paying for a closed frontier service. Confirm the current inference requirements and benchmark protocols before committing a production workflow to it.

Comments

No comments yet. Be the first to comment!