Research
The Research stories that matter — sourced, dated, and explained in plain English, without the hype.
The lead
Latest firstLatest in Research
32-Task Benchmark Shows AI Falls Short on Scientific Figures
32-Task Benchmark Shows AI Falls Short on Scientific Figures: SciDraw-Bench, a 32-task benchmark for scientific figure generation, finds general AI models lag…
SciDraw-Bench Launches to Evaluate AI Scientific Figures
Introduced in a June 24, 2026 arXiv paper, SciDraw-Bench tests 32 tasks across 8 figure types and 10 scientific disciplines to measure…
Al-Mawrid Arabic-English Dictionary IE Achieves High Precision
Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction: A rule-based information extraction method for the…
Recursive Self-Evolving Agents Need Held-Out Selection Gates
New arXiv paper finds recursive self-evolving LLM agents need held-out gates to avoid regressions, as unguarded methods scored 0.14 on WebShop benchmarks.
SciDraw-Bench: 1st AI Scientific Figure Benchmark, 32 Tasks
Released in a June 2026 arXiv preprint, SciDraw-Bench is the first benchmark to evaluate AI-generated scientific figures across 32 tasks, 8 figure…
Miles Connects PyTorch, SGLang, Ray, and Megatron for LLM RL
Miles is an open-source LLM reinforcement-learning stack joining SGLang rollouts, Megatron training, Ray orchestration, and PyTorch extensions.
Coherent Expands Texas InP AI Interconnect Capacity 200%
Coherent Expands Texas InP AI Interconnect Capacity 200%: Coherent is scaling indium phosphide optical component production at its Sherman, Texas campus to…
HP Inc. Scales OpenAI Frontier Partnership for Enterprise AI
HP Inc. is scaling its OpenAI Frontier strategic partnership after successful 2026 pilots, deploying AI across customer support, security, and software…
Hybrid model outperforms transformers on meaning tokens
An Allen AI token-level study shows the 7B Olmo Hybrid predicts content words better than a matched transformer, while the transformer still…
Hybrid LLMs predict meaning tokens better than transformers
AI2 head-to-head testing of 7B Olmo 3 transformer and Olmo Hybrid models finds hybrid LLMs predict meaning tokens better than transformers, with…