Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
Tested on 4 benchmarks with 66.1% average. Top scores: ARC-AGI (89.5%), ARC-AGI-2 (65.1%), Terminal Bench (57.3%).
Qwen3.6 Plus scores 69.8 (99% as good) at $0.33/1M input · 74% cheaper
Complex terminal-based engineering tasks. Models must use command-line tools, navigate filesystems, and debug systems through shell interaction.
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.
ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.
- Typemultimodal
- Context2.0M tokens (~1,000 books)
- ReleasedMar 2026
- LicenseProprietary
- StatusActive
- Cost / Message~$0.005