Scale AI Blog
The Scale Research Team
Jun 2026
Can AI Agents Do the Work of Drug Discovery?
ResearchResearch
May 2026SWE Atlas is Complete: Measuring Coding Agents Across the Engineering Loop
Apr 2026Understanding New Biosecurity Risks Posed by LLMs
Mar 2026Can Coding Agents Become Engineers? We’re Finding Out.
ResearchResearch
Dec 2025MoReBench: Evaluating the Process of AI Moral Reasoning
ResearchResearch
Dec 2025Open-Sourcing MCP-Atlas: A Benchmark for Real Tool Use
Dec 2025Real Speech Breaks AI (And What We're Doing to Fix It)
Nov 2025The Limits of Data Filtering in Bio-Foundation Models
Nov 2025Breaking Out of the Lab: Testing AI in Professional Domains
Oct 2025The Remote Labor Index: Measuring the Automation of Work
ResearchResearch
Oct 2025VisualToolBench: Testing the Limits of AI Vision
Sep 2025SWE-Bench Pro: Raising the Bar for Agentic Coding
Testing & EvalsTesting & Evals
Sep 2025Advancing Agents: Introducing Scale’s Agentic Leaderboards
ResearchResearch
Sep 2025Actions, Not Words: MCP-Atlas Raises the Bar for Agentic Evaluation
Testing & EvalsTesting & Evals
Sep 2025TutorBench: Grading the Next Generation of AI Tutors
ResearchResearch
Sep 2025Using Rubrics to Build Better Models
ResearchResearch
Jul 2025The Future is Multilingual: Scale's New Evaluation Benchmark
ResearchResearch
Jul 2025