Research

Execution is the research problem.

Frontier labs train models to answer. We train models to act: safely, on real infrastructure, with a human in the loop. Measured against the work, not a leaderboard of trivia.

The loop · Perceive, Decide, Act
Perceive
7 modalities

Multimodal capture of real work: screens, documents, SCADA, and the field itself through the Frame SDK.

Decide
calibrated

Confidence-scored reasoning over operations, calibrated against outcomes so the model knows when to defer.

Act
gated

Automated actions behind human-in-the-loop gates. The long arc runs to autonomous wells.

Silver Fox benchmark · v3

Scored on the work, not trivia.

A WCSB execution suite: real field, accounting, and regulatory tasks graded against operator-verified answers. Higher is better.

Apture BasinFrontier · scrubbed promptOpen 13B baseline
Production decline detection+22 pts vs frontier
Apture Basin
94
Frontier · scrubbed
72
Open 13B
61
SCADA alarm triage+12 pts vs frontier
Apture Basin
91
Frontier · scrubbed
79
Open 13B
68
Setpoint drift QA+25 pts vs frontier
Apture Basin
88
Frontier · scrubbed
63
Open 13B
54
Volume reconciliation+15 pts vs frontier
Apture Basin
96
Frontier · scrubbed
81
Open 13B
73
Silver Fox v3 · WCSB execution suite · scored against operator-verified answers. Frontier is a top general model on prompts scrubbed of names, coordinates, and volumes. Numbers are illustrative of the current internal run.
Threads

What we are working on.

All papers →

Read the method. Build on it.

The whitepapers carry the full method and numbers. The SDK and API let you put it to work.