Simulating the
World's Intelligence

Patronus AI is a research lab developing training data and simulated environments for frontier AI.

Building Worlds
That Teach AI

Our research connects a deep investigation of model failures with new methods for developing high-quality tasks and environments for reinforcement learning.

Through scalable world modeling and synthetic data generation, we aim to extend human expertise across more training experiences while preserving the standards of real-world work.

Software Engineering

Building, debugging, and maintaining complex software

Computer Use

Completing tasks across desktop applications, browsers, and tools

Knowledge Work

Synthesizing information, analyzing data, and producing deliverables

Model Alignment

Following instructions, respecting constraints, and exercising sound judgment

Scientific Research

Understanding, synthesizing, and extending scientific research

Simulation Domains

Tasks and environments mirroring work across key industries and functions

Featured Research

Research
World Modeling

Steerable world models for reinforcement learning

Simulating environment dynamics from grounded agent trajectories

Studying transfer to unseen environments

View paper
FigmaTrace

Training data drawn from expert Figma workflows

Preserving design intent when converting workflows into trajectories

Improving agent performance across computer-use environments

View paper
MEMTRACK

Testing long-term memory in dynamic work environments

Tracking changing information across connected platforms

Revealing failures in retaining and applying prior context

View paper
DETOUR

Testing multi-turn search with incomplete information

Resolving ambiguous requests through agent interaction

Spanning text, image, audio, and video

View paper
TRAIL

Investigating failures across complex agent workflows

Human-annotated traces from software and research tasks

Testing models’ ability to locate and explain errors

View paper
TRACE

Detecting reward hacking in coding environments

Comparing intended behavior with exploitative agent trajectories

Testing whether models recognize shortcuts that undermine task goals

View paper
SpeedrunBench

Testing strategy discovery through repeated gameplay

Challenging agents to improve beyond task completion

Measuring planning and execution under time pressure

View paper
FinanceBench

Testing financial reasoning against real company documents

Expert-curated questions with answers and supporting evidence

Revealing limitations in financial question answering

View paper