LifeSciBench
HoldTools
An expert-authored benchmark for evaluating AI on life science research tasks.
Why it's here
Placed in Hold: 1 article(s) of evidence from 1 source(s), led by research-stage coverage, with 0 in the last 30 days. Confidence 24%. Low accumulated evidence, so it defaults conservatively pending more signal.
Evidence (1)
- 6OpenAI Blog·6/17/2026researchOpenAI Introduces LifeSciBench
OpenAI has introduced LifeSciBench, an expert-authored and expert-reviewed benchmark for evaluating how AI systems perform on real-world life science research tasks and decisions. The benchmark is designed to measure model usefulness in practical scientific workflows rather than only abstract test performance.