Research
The Research stories that matter — sourced, dated, and explained in plain English, without the hype.
The lead
Latest firstLatest in Research
CORE-Bench Extends Agent Benchmarks Past Accuracy Saturation
New arXiv research on CORE-Bench v1.1 finds saturated accuracy benchmarks deliver more insight when expanded to measure efficiency, reliability, and ~2x…
OpenAI o3 Model Diagnoses 18 Rare Childhood Genetic Diseases
A June 2026 NEJM AI study finds OpenAI's o3 Deep Research model helped clinicians diagnose 18 previously unsolved rare childhood genetic diseases…
MosaicLeaks study finds AI research agents leak 34% of private enterprise data
New ServiceNow MosaicLeaks research finds unmodified AI research agents leak 34% of private enterprise data via web query logs, with a new…
CS Curriculum Alignment Near-Constant Across CS2013, CS2023
A new arXiv study introduces a reusable human-in-the-loop pipeline to measure CS curriculum alignment against CS2013 and CS2023, finding 49.7% coverage of…
NEJM AI Study: OpenAI o3 Deep Research Adds 4.8% Rare Childhood Disease Yield
AI to help physicians diagnose rare genetic diseases affecting children: A June 2026 NEJM AI peer-reviewed study finds OpenAI o3 Deep Research…
Study of 8 Diffusion Language Models Across 8 Benchmarks Reveals Key Inference Tradeoffs
A June 2026 arXiv preprint analyzing 8 state-of-the-art diffusion language models across 8 benchmarks identifies four high-impact inference design choices and…
OpenAI reasoning model identifies 18 rare childhood genetic diagnoses from 376 unsolved cases
A June 2026 NEJM AI study found an OpenAI o3 reasoning workflow helped clinicians diagnose 18 previously unsolved rare childhood genetic diseases…
OpenAI o3 Deep Research Enables 18 New Rare Childhood Genetic Diagnoses in NEJM AI Study
A June 18 NEJM AI study finds OpenAI's o3 Deep Research reasoning model helped clinicians diagnose 18 previously unsolved rare childhood genetic…
Study: Unstructured Shared Workspace Human-AI Collaboration Cuts Performance Across 1,482 Sessions
A June 2026 arXiv study of 1,482 collaborative sessions finds unstructured shared workspace human-AI collaboration reduces performance, while targeted HITL…
MosaicLeaks benchmark finds deep research agents leak private data in 34% of test chains
New ServiceNow MosaicLeaks benchmark finds standard deep research agents leak private enterprise data in 34% of multi-hop test chains, with performance-focused…