engblogs

summaries of the latest blog articles from your favorite tech companies.
OpenAIOpenAI

Separating signal from noise in coding evaluations

A detailed audit reveals pervasive evaluation flaws in SWE-Bench Pro and related benchmarks, arguing for stronger signal, longer-horizon tasks, improved data quality checks, human-in-the-loop review, and cautious deployment decisions under OpenAI’s Preparedness Framework.

7/8/2026
OpenAIOpenAI

Introducing GPT-Live

GPT-Live is a continuous, full-duplex voice system that talks and listens in real time, delegates complex tasks to frontier models like GPT-5.5, and integrates provenance signals and watermarking with a verification API to enable safer, more natural ChatGPT Voice experiences.

7/8/2026
OpenAIOpenAI

Our approach to government and national security partnerships

OpenAI articulates National Security Principles to guide government partnerships and frontier AI use in national security, emphasizing transparency, democratic accountability, meaningful human judgment, and safeguards across cyber defense and biosecurity contexts.

7/8/2026
OpenAIOpenAI

Helping K–12 educators build practical AI skills

OpenAI Academy's AI Skills Jam for K-12 educators is a hands-on, in-person program with mentor support that helps teachers and district leaders translate AI into practical, scalable improvements in teaching, planning, and communication.

7/8/2026
OpenAIOpenAI

Separating signal from noise in coding evaluations

A rigorous audit of SWE-Bench Pro and its predecessor reveals widespread evaluation flaws and demonstrates how to separate signal from noise in coding benchmarks, outlining a QA-driven pipeline that leverages both automated and human reviews to ensure fair, meaningful model capability assessments.

7/8/2026
OpenAIOpenAI

Our approach to government and national security partnerships

OpenAI's National Security Principles outline responsible government and national security partnerships for frontier AI, emphasizing transparency, democratic accountability, and safeguards while enabling cyber defense and biosecurity applications.

7/8/2026
Snorkel AISnorkel AI

Grok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work

Grok 4.5 delivers the strongest mean performance on ~2,000 GDPval+ professional tasks, outperforming GPT 5.5 and Opus 4.8 across domains with lower error rates and expert-rubric satisfaction.

7/8/2026
Snorkel AISnorkel AI

Grok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work

Grok 4.5 delivers the strongest mean pass rate on Snorkel’s GDPval+ professional-work tasks, outperforming GPT 5.5 and Opus 4.8 across domains like legal, education, healthcare, and QA while showing lower error rates and more robust failure handling.

7/8/2026
GitHubGitHub

Automating cross-repo documentation with GitHub Agentic Workflows

A case study on automating cross-repo documentation with GitHub Agentic Workflows, mapping product milestones to release branches and delivering SME-reviewed docs via secure, token-scoped pull requests in a separate docs repo.

7/8/2026
GitHubGitHub

Automating cross-repo documentation with GitHub Agentic Workflows

Automates cross-repo, SME-reviewed documentation by generating PRs from feature changes via GitHub Agentic Workflows, using a constrained per-workflow bot and a safe-outputs pipeline to securely publish docs across repos.

7/8/2026
OpenAIOpenAI

MUFG aims to become AI-native with OpenAI

MUFG partners with OpenAI to roll out ChatGPT Enterprise across Mitsubishi UFJ Bank, accelerating an AI-native transformation with enterprise-wide adoption and AI-powered customer experiences in a secure, governed digital bank future.

7/7/2026
OpenAIOpenAI

Australian Payments Plus moves faster with ChatGPT and Codex

AP+ accelerates product development, investigations, and member communications by deploying ChatGPT Enterprise and Codex to speed complex issue resolution, simulate payment journeys, and strengthen governance in a regulated payments ecosystem.

7/7/2026