Results

Same events. Who grades better?

We gave two AI models the exact same events ARGUS's formulas saw, then checked who graded them better. The models are good at sorting. They slip on the calls that matter most: how risky it is, whether a claim holds up, and how sure to be.

Verdict

Use a model to draft or sort. Use ARGUS for the grades that decide what enters the brief.

Job Question What we saw Better
Spot a surge Is activity jumping? Models matched 100% Either
Rank places Which place looks hotter? Models ranked well Models (sorting)
Name the risk LOW or CRITICAL? ~11% matched ARGUS ARGUS
Test a claim Support or contradict? ~40–53% matched ARGUS
Say how sure How confident? GPT overclaimed ARGUS

How we tested: 14 events across 3 fixed packs, run through two AI models (Claude Haiku 4.5 and GPT-4o-mini, temperature 0) and ARGUS's formulas on the same inputs. A small, deliberate set: every event, prompt, and score is in the repo so you can re-run it yourself.

Formulas: Methods · Full replication under research/llm-eval/.

Get ARGUS

GitHub · Guide · About