Compare · ModelsLive · 2 picked · head to head
Phi-1.5 vs RedPajama-INCITE-7B-Base
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Phi-1.5 wins on 8/11 benchmarks
Phi-1.5 wins 8 of 11 shared benchmarks. Leads in knowledge · general · math.
Category leads
knowledge·Phi-1.5general·Phi-1.5language·RedPajama-INCITE-7B-Basemath·Phi-1.5reasoning·Phi-1.5
Hype vs Reality
Attention vs performance
Phi-1.5
#260 by perf·no signal
RedPajama-INCITE-7B-Base
#257 by perf·no signal
Vendor risk
Who is behind the model
Microsoft
$3.00T·Big Tech
U
Unknown
private · undisclosed
Head to head
11 benchmarks · 2 models
Phi-1.5RedPajama-INCITE-7B-Base
ARC AI2
Phi-1.5 leads by +7.1
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
Phi-1.5
25.9
RedPajama-INCITE-7B-Base
18.8
HellaSwag
RedPajama-INCITE-7B-Base leads by +30.3
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
Phi-1.5
30.1
RedPajama-INCITE-7B-Base
60.4
BBH (HuggingFace)
Phi-1.5 leads by +2.4
Phi-1.5
7.5
RedPajama-INCITE-7B-Base
5.1
GPQA
Phi-1.5 leads by +1.7
Phi-1.5
2.4
RedPajama-INCITE-7B-Base
0.7
IFEval
RedPajama-INCITE-7B-Base leads by +0.5
Phi-1.5
20.3
RedPajama-INCITE-7B-Base
20.8
MATH Level 5
Phi-1.5 leads by +0.2
Phi-1.5
1.8
RedPajama-INCITE-7B-Base
1.6
MMLU-PRO
Phi-1.5 leads by +5.5
Phi-1.5
7.7
RedPajama-INCITE-7B-Base
2.2
MUSR
Phi-1.5 leads by +0.4
Phi-1.5
3.4
RedPajama-INCITE-7B-Base
3.0
MMLU
Phi-1.5 leads by +15.1
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Phi-1.5
16.8
RedPajama-INCITE-7B-Base
1.7
OpenBookQA
RedPajama-INCITE-7B-Base leads by +3.7
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
Phi-1.5
16.3
RedPajama-INCITE-7B-Base
20.0
Winogrande
Phi-1.5 leads by +19.2
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
Phi-1.5
46.8
RedPajama-INCITE-7B-Base
27.6
Full benchmark table
| Benchmark | Phi-1.5 | RedPajama-INCITE-7B-Base |
|---|---|---|
ARC AI2 AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval. | 25.9 | 18.8 |
HellaSwag HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios. | 30.1 | 60.4 |
BBH (HuggingFace) | 7.5 | 5.1 |
GPQA | 2.4 | 0.7 |
IFEval | 20.3 | 20.8 |
MATH Level 5 | 1.8 | 1.6 |
MMLU-PRO | 7.7 | 2.2 |
MUSR | 3.4 | 3.0 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 16.8 | 1.7 |
OpenBookQA OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting. | 16.3 | 20.0 |
Winogrande WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs. | 46.8 | 27.6 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
U RedPajama-INCITE-7B-Base | — | — | — | — |