data for frontier models deploying into open‑ended environments

We create training-ready datasets using real open-ended workflows in dynamic environments that scripted tasks can't replicate. Custom datasets and environments available.

Desktop activity decomposed into workspace state, actions, and workflow timelines.
600+
hours of trajectories
350k+
annotated action steps
27k+
app transitions
46+
transitions per hour
45+
unique apps covered

first in class learning signals

interwoven workflows

Multiple independent tasks being executed together using the shared context on screen. Natural task sequencing, interruption, resumption, and completion with visible cues like notifications and load states.

full workspace state

Entire desktop provenance with multi-monitor support spanning a variety of resolutions, windows, and contexts. Demonstrated information transformations across app and window boundaries.

open-ended workflows

Unscripted multi-thousand-step workflows spanning tens of hours without synthetic data, scripted instructions, or app clones.

trajectory viewer

Browsers
Brave Browser, Google Chrome, Dia, Firefox, Chromium
Communication
Slack, Mail, Superhuman, Messages, Discord, FaceTime
Coding & AI
Cursor, Terminal, Ghostty, Xcode, Electron, Claude, ChatGPT, Zo, Grok Bot
Knowledge & docs
Notion, Obsidian, Notes, TextEdit, Pages, Granola, Numbers, Preview, Dictionary
Design & media
Figma, CleanShot X, Screenshot, QuickTime Player, Photos, Spotify
Desktop & planning
Finder, Calendar, Notion Calendar, Todoist, Sunsama, Opal, Calculator, System Settings, Font Book, Archive Utility, Maps
Recording
H.264 video at 30 fps, with multi-monitor support
Input alignment
Frame-aligned raw input events with sub-millisecond timestamps
Step-level actions
Discrete step-level actions reduced from raw input events in JSON
Window context
Active application and window metadata derived from the accessibility tree
State
Accessibility trees, pruned HTML, and DOM snapshots
Task Segmentation
Decomposition of trajectories into segmented workflows and atomic tasks

faq

Can the corpus be expanded to meet my training volume?

Yes. We can already generate thousands of hours of data, and are scaling our infrastructure and tooling toward tens of thousands of hours per month. If your team needs a larger volume, please reach out.

Can Fig build custom datasets?

Yes. We build custom data around your requirements, including new environments, traces, and trajectories across the applications, domains, and workflows you care about. We're already working with teams building benchmarks, evals, models, robotics, and manufacturing. Tell us what you need and we'll scope it with you.

How is this different from other computer-use datasets?

Fig's datasets bring first in class learning signals that cannot exist in scripted tasks. Through deep observational telemetry across real work as it unfolds, the corpus captures real-world, open-ended workflows executed in continuously evolving environments. This provides customers with data that represents the exact type of environments and operations their models will be deployed to in production, which scripted collection can't reproduce.

Who creates Fig's data?

Every trajectory is captured using Fig's proprietary capture tooling from consenting operators completing real work across research, operations, and engineering. There are no assigned tasks, prompts, or scripts, each recording contains the full state of the computer as the work happened, interruptions included.

Can I use this data to train models, build evals, or create benchmarks?

Yes. Our data is built to be used by frontier teams across a wide variety of training applications. We're happy to walk through how this data or one of our custom solutions can fit into your existing pipeline. Tell us what you're building and we'll go from there.

explore custom solutions.

We work with frontier teams to build custom solutions tailored to their model-development requirements. Let’s explore how we can collaborate.
get started