OpenAIRunning Codex safely at OpenAI
Running Codex safely at OpenAI: deploying bounded, sandboxed execution with approval workflows and agent-aware telemetry to balance developer productivity with enterprise security and auditable governance.
OpenAIRunning Codex safely at OpenAI
Codex is engineered for safe, enterprise deployment by enforcing bounded execution, sandboxed actions, approval gates, strict network and credential controls, and agent-native telemetry for auditability.
OpenAIRunning Codex safely at OpenAI
Deploys Codex within strict sandboxed boundaries, uses approval workflows, and leverages agent-aware telemetry and compliance logging to balance rapid development with enterprise security.
PinterestEnhancing Ad Relevance: Integrating Real-Time Context into Sequential Recommender Models
A Contextual Sequential Two-Tower Model that injects real-time context into a Transformer-based recommender, trained with synthetic context data and deployed via a hybrid offline-online serving flow to boost Related Pins ad relevance and ROAS.
PinterestEnhancing Ad Relevance: Integrating Real-Time Context into Sequential Recommender Models
Introducing a Contextual Sequential Two-Tower model that injects real-time context through a context layer and synthetic data, enabling hybrid offline-online inference to boost ad relevance on Related Pins.
Berkeley AIAdaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Adaptive Parallel Reasoning (APR) enables models to dynamically balance parallel and sequential reasoning at inference time, using fork-join architectures and KV-cache strategies to improve latency and throughput across engine-modified and engine-agnostic implementations.
Berkeley AIAdaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
Explores adaptive parallel reasoning as a paradigm shift for inference, detailing how models can dynamically balance sequential and parallel execution, orchestrate multiple threads, and manage KV cache to achieve faster, more efficient reasoning with flexible, engine-agnostic design.
Modular AIModular: Why LLM Inference Needs a New Kind of Router - Part 1
Explores how Modular Cloud reimagines LLM inference routing with cache-aware, multi-step orchestration and disaggregated GPU pods to achieve predictable latency in stateful backends.
OpenAISimplex rethinks software development with Codex
Simplex rethinks software delivery by deploying Codex with ChatGPT Enterprise to accelerate design, development, and testing with measurable productivity gains.
OpenAIIntroducing Trusted Contact in ChatGPT
Overview of Trusted Contact in ChatGPT, a safety feature enabling a trusted adult to be notified during potential self-harm crises to connect users with real-world support while preserving privacy and undergoing human review.
OpenAIAdvancing voice intelligence with new models in the API
Introducing three real-time voice models—GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper—that enable live reasoning, multilingual translation, and streaming transcription to power end-to-end voice interfaces that listen, understand context, translate in real time, transcribe as you speak, and take action within conversations.
OpenAISimplex rethinks software development with Codex
Simplex rethinks software delivery by deploying ChatGPT Enterprise and Codex to validate AI-driven development and accelerate productivity, achieving 70% faster screen development, 40% faster design, and 17% faster internal integration testing through AI-native workflows.