OpenAISeparating signal from noise in coding evaluations
A detailed audit reveals pervasive evaluation flaws in SWE-Bench Pro and related benchmarks, arguing for stronger signal, longer-horizon tasks, improved data quality checks, human-in-the-loop review, and cautious deployment decisions under OpenAI’s Preparedness Framework.
OpenAIIntroducing GPT-Live
GPT-Live is a continuous, full-duplex voice system that talks and listens in real time, delegates complex tasks to frontier models like GPT-5.5, and integrates provenance signals and watermarking with a verification API to enable safer, more natural ChatGPT Voice experiences.
OpenAIOur approach to government and national security partnerships
OpenAI articulates National Security Principles to guide government partnerships and frontier AI use in national security, emphasizing transparency, democratic accountability, meaningful human judgment, and safeguards across cyber defense and biosecurity contexts.
OpenAIHelping K–12 educators build practical AI skills
OpenAI Academy's AI Skills Jam for K-12 educators is a hands-on, in-person program with mentor support that helps teachers and district leaders translate AI into practical, scalable improvements in teaching, planning, and communication.
OpenAISeparating signal from noise in coding evaluations
A rigorous audit of SWE-Bench Pro and its predecessor reveals widespread evaluation flaws and demonstrates how to separate signal from noise in coding benchmarks, outlining a QA-driven pipeline that leverages both automated and human reviews to ensure fair, meaningful model capability assessments.
OpenAIOur approach to government and national security partnerships
OpenAI's National Security Principles outline responsible government and national security partnerships for frontier AI, emphasizing transparency, democratic accountability, and safeguards while enabling cyber defense and biosecurity applications.
Snorkel AIGrok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work
Grok 4.5 delivers the strongest mean performance on ~2,000 GDPval+ professional tasks, outperforming GPT 5.5 and Opus 4.8 across domains with lower error rates and expert-rubric satisfaction.
Snorkel AIGrok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work
Grok 4.5 delivers the strongest mean pass rate on Snorkel’s GDPval+ professional-work tasks, outperforming GPT 5.5 and Opus 4.8 across domains like legal, education, healthcare, and QA while showing lower error rates and more robust failure handling.
GitHubAutomating cross-repo documentation with GitHub Agentic Workflows
A case study on automating cross-repo documentation with GitHub Agentic Workflows, mapping product milestones to release branches and delivering SME-reviewed docs via secure, token-scoped pull requests in a separate docs repo.
GitHubAutomating cross-repo documentation with GitHub Agentic Workflows
Automates cross-repo, SME-reviewed documentation by generating PRs from feature changes via GitHub Agentic Workflows, using a constrained per-workflow bot and a safe-outputs pipeline to securely publish docs across repos.
OpenAIMUFG aims to become AI-native with OpenAI
MUFG partners with OpenAI to roll out ChatGPT Enterprise across Mitsubishi UFJ Bank, accelerating an AI-native transformation with enterprise-wide adoption and AI-powered customer experiences in a secure, governed digital bank future.
OpenAIAustralian Payments Plus moves faster with ChatGPT and Codex
AP+ accelerates product development, investigations, and member communications by deploying ChatGPT Enterprise and Codex to speed complex issue resolution, simulate payment journeys, and strengthen governance in a regulated payments ecosystem.