AIWARLIVE
← Command board

Llama — Field Evidence

paperUnverified
Llama · cs.CL · praise
Field note
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness — Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-sea
collected 2026-07-22original 2026-07-21

Does this shift the US–China race?

Be the first to call it

Impact on the index
Benchmark win · minor
US +30China -7
Directional contribution — before recency decay and per-type diminishing returns. How it works →
Source ↗

Related Front

US vs ChinaLikely
U.S. FrontiervsChina Open-Weight

Likely — capability gap narrowing on common tasks

Letters from the Front

Be the first to file a report.