Research Hub · sourced & dated

Research

The Research stories that matter — sourced, dated, and explained in plain English, without the hype.

22 stories published 100% sourced & dated

The lead

Latest first

Latest in Research

AI CORE-Bench Extends Agent Benchmarks Past Accuracy Saturation

CORE-Bench Extends Agent Benchmarks Past Accuracy Saturation

New arXiv research on CORE-Bench v1.1 finds saturated accuracy benchmarks deliver more insight when expanded to measure efficiency, reliability, and ~2x…

Aira · Jun 28, 2026
AI OpenAI o3 Model Diagnoses 18 Rare Childhood Genetic Diseases

OpenAI o3 Model Diagnoses 18 Rare Childhood Genetic Diseases

A June 2026 NEJM AI study finds OpenAI's o3 Deep Research model helped clinicians diagnose 18 previously unsolved rare childhood genetic diseases…

Aira · Jun 21, 2026
AI MosaicLeaks study finds AI research agents leak 34% of private enterprise data

MosaicLeaks study finds AI research agents leak 34% of private enterprise data

New ServiceNow MosaicLeaks research finds unmodified AI research agents leak 34% of private enterprise data via web query logs, with a new…

Aira · Jun 21, 2026
AI CS Curriculum Alignment Near-Constant Across CS2013, CS2023

CS Curriculum Alignment Near-Constant Across CS2013, CS2023

A new arXiv study introduces a reusable human-in-the-loop pipeline to measure CS curriculum alignment against CS2013 and CS2023, finding 49.7% coverage of…

Aira · Jun 20, 2026
AI NEJM AI Study: OpenAI o3 Deep Research Adds 4.8% Rare Childhood Disease Yield

NEJM AI Study: OpenAI o3 Deep Research Adds 4.8% Rare Childhood Disease Yield

AI to help physicians diagnose rare genetic diseases affecting children: A June 2026 NEJM AI peer-reviewed study finds OpenAI o3 Deep Research…

Aira · Jun 20, 2026
AI Study of 8 Diffusion Language Models Across 8 Benchmarks Reveals Key Inference Tradeoffs

Study of 8 Diffusion Language Models Across 8 Benchmarks Reveals Key Inference Tradeoffs

A June 2026 arXiv preprint analyzing 8 state-of-the-art diffusion language models across 8 benchmarks identifies four high-impact inference design choices and…

Aira · Jun 20, 2026
AI OpenAI reasoning model identifies 18 rare childhood genetic diagnoses from 376 unsolved cases

OpenAI reasoning model identifies 18 rare childhood genetic diagnoses from 376 unsolved cases

A June 2026 NEJM AI study found an OpenAI o3 reasoning workflow helped clinicians diagnose 18 previously unsolved rare childhood genetic diseases…

Aira · Jun 20, 2026
AI OpenAI o3 Deep Research Enables 18 New Rare Childhood Genetic Diagnoses in NEJM AI Study

OpenAI o3 Deep Research Enables 18 New Rare Childhood Genetic Diagnoses in NEJM AI Study

A June 18 NEJM AI study finds OpenAI's o3 Deep Research reasoning model helped clinicians diagnose 18 previously unsolved rare childhood genetic…

Aira · Jun 19, 2026
AI Study: Unstructured Shared Workspace Human-AI Collaboration Cuts Performance Across 1,482 Sessions

Study: Unstructured Shared Workspace Human-AI Collaboration Cuts Performance Across 1,482 Sessions

A June 2026 arXiv study of 1,482 collaborative sessions finds unstructured shared workspace human-AI collaboration reduces performance, while targeted HITL…

Aira · Jun 19, 2026
AI MosaicLeaks benchmark finds deep research agents leak private data in 34% of test chains

MosaicLeaks benchmark finds deep research agents leak private data in 34% of test chains

New ServiceNow MosaicLeaks benchmark finds standard deep research agents leak private enterprise data in 34% of multi-hop test chains, with performance-focused…

Aira · Jun 18, 2026
The zBrandco Edition

Never miss what matters in Research.