Results
Same events. Who grades better?
We gave two AI models the exact same events ARGUS's formulas saw, then checked who graded them better. The models are good at sorting. They slip on the calls that matter most: how risky it is, whether a claim holds up, and how sure to be.
Verdict
Use a model to draft or sort. Use ARGUS for the grades that decide what enters the brief.
| Job | Question | What we saw | Better |
|---|---|---|---|
| Spot a surge | Is activity jumping? | Models matched 100% | Either |
| Rank places | Which place looks hotter? | Models ranked well | Models (sorting) |
| Name the risk | LOW or CRITICAL? | ~11% matched ARGUS | ARGUS |
| Test a claim | Support or contradict? | ~40–53% matched | ARGUS |
| Say how sure | How confident? | GPT overclaimed | ARGUS |
How we tested: 14 events across 3 fixed packs, run through two AI models (Claude Haiku 4.5 and GPT-4o-mini, temperature 0) and ARGUS's formulas on the same inputs. A small, deliberate set: every event, prompt, and score is in the repo so you can re-run it yourself.
Formulas: Methods
·
Full replication under research/llm-eval/.