SoundCloudLess Is More: Why Audio on SoundCloud Looks Different
Explains why SoundCloud's Fraunhofer libfdk_aac upgrade appears like a downgrade—cutting the top frequencies to save bits—and how this tradeoff improves overall perceptual quality, shown with spectrograms and blind tests.
Snorkel AIJudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
JudgmentBench provides an empirical comparison of rubric-based scoring and comparative judgment for quality assessment in knowledge-work domains, showing that pairwise preference judgments nearly perfectly recover quality orderings (Spearman's rho ≈ 0.908) and require about half the annotation time, outperforming rubrics on BigLawBench tasks.
Snorkel AIJudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
JudgmentBench pits comparative judgment against rubric-based scoring on 30 BigLawBench tasks to show expert pairwise preferences better recover quality ordering and cut annotation time, informing AI benchmarking in legal knowledge work.
MIT AIMIT in the media: For the future of tech, "Massachusetts can absolutely lead"
MIT reinforces Massachusetts’ tech leadership by advancing AI applications, entrepreneurship and startup culture, and cross-disciplinary research in energy, quantum, and defense technology, supported by industry partnerships and the MIT ecosystem.
MIT AIMIT in the media: For the future of tech, "Massachusetts can absolutely lead"
Explores MIT's role as a hub for AI, entrepreneurship, and groundbreaking energy and quantum research, spotlighting Massachusetts as a global tech leader.
Modular AIModular: Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More
Modular 26.4 accelerates SOTA MoE serving and MAX-enabled model bring-up with agentic skills, adds Mojo 1.0 beta 2 stabilization, broader model and hardware support via Modular Cloud, and enhanced OpenAI API compatibility.
OpenAIIntroducing LifeSciBench
LifeSciBench: a comprehensive, expert-authored benchmark that evaluates AI systems on real-world life science research tasks across seven workflows and seven biological domains, using 1,062 task artifacts and rigorous rubrics to assess scientific reasoning, uncertainty handling, and translational potential.
OpenAIA near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
GPT‑5.4 powered Maria autonomously proposes and validates a TEMPO‑assisted improvement to Chan–Lam coupling in medicinal chemistry via a high‑throughput lab, with human experts overseeing the process.
OpenAIIntroducing LifeSciBench
Introducing LifeSciBench: a rigorously designed, expert-curated benchmark that measures whether Agentic AI can meaningfully support real-world life-science research across seven workflows and seven biological domains, using artifact-rich tasks and granular rubrics.
OpenAIA near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
A near-autonomous AI chemist integrated with a high-throughput lab uses GPT-5.4 to identify TEMPO as an additive that markedly improves Chan-Lam coupling yields across diverse substrates, showcasing scalable AI-assisted medicinal chemistry with human oversight.
OpenAIIntroducing LifeSciBench
LifeSciBench is a comprehensive, expert-authored benchmark that tests AI systems on realistic life-science research tasks—spanning seven workflows and seven domains, with artifact-rich prompts, rigorous rubrics, and a focus on evidence handling and uncertainty in drug-discovery settings.
OpenAIA near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
A near-autonomous AI chemist, Maria, working with Molecule.one and GPT‑5.4 in a high-throughput lab, boosts Chan–Lam coupling yields using TEMPO across diverse substrates, illustrating AI-assisted discovery with human oversight in medicinal chemistry.