Scale partners with Mayo Clinic to develop reliable AI for healthcareRead the Full Story

Scale at ICML 2026

90% of the world's leading generative AI model builders are powered by Scale.

Don't Miss Our Sessions at ICML

Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections

Speakers
  • Brad KenstlerDirector, Agent Capabilities and Environments
Time

Wed, Jul 8, 2026 • 2:30 PM – 4:15 PM KST · HALL A #703

VeRO: An Evaluation Harness for Agents to Optimize Agents

Speakers
  • Varun UrsekarMachine Learning Research Engineer
  • Apaar ShankerSr. Machine Learning Research Engineer
  • Veronica ChatrathMachine Learning Product Manager
Time

Tue, Jul 7, 2026 • 2:00 PM – 3:45 PM KST · HALL A #1810

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

Speakers
  • Andrew KlearmanApplied AI Engineer
Time

Tue, Jul 7, 2026 • 2:00 PM – 3:45 PM KST · HALL A #4307

Online Rubrics Elicitation from Pairwise Comparisons

Speakers
  • Yunzhong HeHead of Post-Training & Eval Research
  • Feyza AkyurekResearch Scientist
Time

Thu, Jul 9, 2026 • 10:30 AM – 12:15 PM KST · HALL A #114

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Speakers
  • Yannis HeAI Product Manager, Scale AI
  • Miguel Romero Calvo
  • Brad KenstlerDirector, Agent Capabilities and Environments
Time

Thu, Jul 9, 2026 • 10:30 AM – 12:15 PM KST · HALL A #4624

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

Speakers
  • Madhu SehwagSecurity and Policy Research Lead
Time

Thu, Jul 9, 2026 • 2:30 PM – 4:15 PM KST · HALL A #1500

RubricRobustness: A Simple Framework for Evaluating the Robustness of Rubrics-Based Benchmarks

Speakers
  • Manasi Sharma
Time

Thu, Jul 9, 2026 • 10:30 AM – 12:15 PM KST · HALL A #3605

ROK-FORTRSS

Speakers
  • Madhu Sehwag
Time

Workshop: Thu, Jul 9, 2026 • 4:00 PM – 1:00 AM PDT · ROOM 317

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate

Speakers
  • Rodney Lafuente-Mercado
Time

Workshop: Fri, Jul 10, 2026 • 4:00 PM – 1:00 AM PDT · ASEM Ballroom 203

Scale Data Engine

Trusted by the world's leading ML teams to accelerate model development

Quality

Scale can provide the core tenet of any dataset with high-quality labels from domain experts.

Cost Effective

Easily find, categorize, and fix model failures with Scale's Data Engine. Then, optimize labeling spend with high-value curated data.

Scalability

Scale's data engine can support any ML project from lower-volume experiments to high-volume production projects. Scale up, or down, as needed.

Diversity

Scale delivers the greatest variety and diversity of data to help deliver the greatest value to your model performance.

Enabling the Next Generation of LLMs

Ops Center for Quality Control

Real-time visibility into data collection and curation.

Experts, Linguists, and Coders

Access a global network of hand-picked experts across diverse fields to build the highest quality datasets.

Improved Models

Train your models with advanced datasets delivered through our purpose-built infrastructure.

Increased Efficiency

Faster, more cost-effective dataset creation.

Model Evaluation

Scale proactively finds and surfaces model weaknesses, including targeted red-teaming.

Responsible Development

Upholding privacy, fairness, transparency and ethics.

We set the benchmark for what’s possible with AI