GPT — Field Evidence
paperUnverified
GPT · GPT · complaint
Field note
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs — Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety g
collected 2026-07-21original 2026-07-20
Does this shift the US–China race?
Be the first to call it
Impact on the index
Usage limits · minor
US -12
Directional contribution — before recency decay and per-type diminishing returns. How it works →
Related Front
US vs ChinaLikely
U.S. FrontiervsChina Open-Weight
Likely — capability gap narrowing on common tasks
Letters from the Front
Be the first to file a report.