OpenAIHow ChatGPT adoption has expanded
Global expansion of ChatGPT adoption is revealed by OpenAI Signals, showing deeper usage, more diverse tasks, and increasing non-English engagement across regions.
OpenAIInside Genebench-Pro
Inside Genebench-Pro offers a concise, technical tour of the benchmark's design, datasets, prompts, and ten case studies—from lncRNA dependencies to cis-MVMR disease-effect estimation and Hi-C loop analysis—highlighting supporting materials and evaluation insights.
OpenAIIntroducing GeneBench-Pro
GeneBench-Pro redefines AI-driven evaluation by offering a synthetic, research-level benchmark that tests higher-level reasoning, judgment, and decision-making in computational biology across genomics and translational medicine.
OpenAICore dump epidemiology: fixing an 18-year-old bug
An epidemiological, population-based debugging journey uncovers two independent crash populations in OpenAI's Rockset data infrastructure—a hardware-induced misalignment and an 18-year-old GNU libunwind race—solved by building a high-quality core-dump data set and reframing debugging from single cases to population analytics.
OpenAIInside Genebench-Pro
A technical tour of Genebench-Pro's benchmark suite, detailing prompts, datasets, and supporting materials across TXR1 inhibitors, lncRNA locus studies, cis-MVMR disease effects, single-cell RNA-seq correction, Hi-C loop analysis, and ancestry reconstruction pipelines.
OpenAIHow ChatGPT adoption has expanded
A data-driven look at how ChatGPT adoption is expanding globally, with deeper user engagement, more diverse tasks, and rising non-English usage across regions.
OpenAICore dump epidemiology: fixing an 18-year-old bug
Population-scale debugging reveals how an 18-year-old GNU libunwind race condition—unmasked by core dumps and hardware quirks—caused crashes in OpenAI's Rockset data infrastructure and how a data-driven epidemiology approach fixed it.
OpenAIHow ChatGPT adoption has expanded
A concise, data-driven look at how ChatGPT adoption is expanding globally, driven by OpenAI Signals insights into deeper usage, multilingual engagement, and accessible plans.
Snorkel AIAgents’ Last Exam: AI Benchmarking for Real Work
A technical blogpost exploring ALE, the living benchmark that tests AI agents on long-horizon, real-world workflows with verifiable outcomes to assess deployment readiness and economic impact.
Snorkel AIAgents’ Last Exam: AI Benchmarking for Real Work
Examines Agents' Last Exam (ALE), a living, expert-verified benchmark that measures AI agents on long-horizon, real-world professional workflows with verifiable outcomes across 55 fields, highlighting Generalist Computer-Use Agents (GCUA) and the gap between frontier models and deployment readiness.
MIT AIQ&A: What is agentic AI today, and what do we want it to be?
A technical overview of agentic AI, detailing how action-taking agents differ from generative models, how wrappers and tools extend foundation models, the training-data challenges, the rise of coding agents, and the future of multimodal architectures to enable real-world interaction.
MIT AIQ&A: What is agentic AI today, and what do we want it to be?
A technical overview of agentic AI today—how agents use tools and wrappers to take actions in the world, how this differs from generative models, notable applications and risks, and possible future directions.