Where AI learns to do real work

Labelbox builds the environments frontier labs use to train AI and the platform enterprises use to put agents to work.

Recursion· Sessions
An acquisition’s due diligence, worked end to end.
A coordinator and five specialists read the data room in Box, check the courts and the audit, then draft the investment committee memo for review in Slack. A grader agent scores the outcome.

For enterprises

Agents for the work that keeps coming back

Recursion isn’t a chatbot. Its agents start on their own, on a schedule, an alert or an event, and work in the background until the job is done: every ready ticket built, every new bug fixed, every dependency kept current, dozens of experiments run at once. Here are a few of the jobs teams hand it.

  • Ticket moved to ReadyLinearGitHubVercelSlack
    Delivers: A tested pull request with a preview link, ready for review
    Engineering

    Feature Builder

    Builds each ready ticket into a working feature, with tests, in its own sandbox.

  • New issue in SentrySentryGitHubLinear
    Delivers: A pull request with the fix and a test that proves it
    Engineering

    Bug Fixer

    Reproduces each new bug, finds the cause and writes the fix with a regression test.

  • New package version releasedGitHubJenkinsSlack
    Delivers: One ready-to-merge pull request per package, tests green
    Engineering

    Dependency Upgrader

    Upgrades dependencies as new versions ship, fixes what breaks and reruns every test.

  • PagerDuty alertPagerDutyDatadogSentryGitHub
    Delivers: A root-cause draft and the suspect commit before on-call is up
    Engineering

    Incident Investigation

    Correlates alerts, logs and recent deploys the moment an alert fires.

  • CI failure on mainGitHubJenkinsSlack
    Delivers: A pull request with the fix and 50 green runs in a row
    Engineering

    Flaky Test Fixer

    Reproduces each flaky test, finds the race and fixes it.

  • Spec approved in NotionNotionGitHubVercel
    Delivers: A live prototype on a preview link, with the spec beside it
    Product

    Prototype Builder

    Turns each approved spec into a working prototype your team can click through.

  • Request tagged for the roadmapLinearGongNotion
    Delivers: A draft spec with the problem, the customer evidence and open questions
    Product

    Spec Writer

    Writes the first spec for each request that makes the roadmap, from the calls and tickets behind it.

  • Pull request mergedGitHubLinearSlack
    Delivers: A docs pull request that matches what shipped
    Product

    Docs Updater

    Updates your docs whenever a merged pull request changes how the product works.

  • Feature released in LinearLinearSalesforceGmail
    Delivers: A personal note to each requester, queued for the account owner
    Product

    Close the Loop

    When a requested feature ships, writes to every customer who asked for it.

  • Mondays 07:00GongSalesforceNotion
    Delivers: Requests ranked by the customers and revenue behind them
    Product

    Feedback Digest

    Reads the week’s sales calls, reviews and feature requests for what customers keep asking for.

  • Issue labeled sweepLinearGitHubDatabricks
    Delivers: A results table with the winning config, rerun to confirm
    Research

    Experiment Sweep

    Runs dozens of experiment configs at once and compares every run with the baseline.

  • Pull request touching promptsGitHubGoogle BigQuerySlack
    Delivers: A per-task score diff, and a blocked merge if anything regresses
    Research

    Eval Regression Watch

    Reruns your eval suite on every model or prompt change.

  • New dataset versionDatabricksGitHubSlack
    Delivers: A checkpoint, its eval report and a model card
    Research

    Model Trainer

    Launches fine-tuning runs, watches the curves and stops bad runs early.

  • New model releasedDatabricksGoogle BigQuerySlack
    Delivers: A leaderboard on your tasks, with cost and latency for each model
    Research

    Benchmark Runner

    Runs every candidate model against your own benchmarks in parallel, overnight.

  • Mondays 09:00NotionGitHubSlack
    Delivers: A digest of what reproduced, with the code and numbers
    Research

    Research Digest

    Reads new papers and repos on your topics and reruns the key result.

  • Alert webhookOktaGoogle Cloud LoggingSlack
    Delivers: Every alert investigated, with a recommended containment step
    Security

    Threat Investigation

    Investigates every identity, mailbox and endpoint alert as it fires.

  • New pull requestGitHubJira
    Delivers: Real vulnerabilities only, each with proof and the fix
    Security

    Code Security Review

    Reviews every pull request and reproduces what it finds.

  • Critical CVE publishedGitHubJiraSlack
    Delivers: A tested patch for every exposed service, most urgent first
    Security

    Vulnerability Patcher

    Finds every service a new critical vulnerability touches and patches it.

  • Nightly 02:00Google Cloud LoggingGitHubJira
    Delivers: Each real exposure, ranked, with the fix as a pull request
    Security

    Cloud Posture Review

    Checks every cloud account for misconfigurations and exposed data.

  • First day of each quarterOktaGitHubGoogle Drive
    Delivers: Every stale or risky grant, with a revoke list for approval
    Security

    Access Review

    Reviews who can reach what, and flags access no one uses or should have.

  • Mondays 06:00SalesforceSnowflakeGoogle Drive
    Delivers: An updated forecast with what moved since last week, and why
    Finance

    Forecast Refresh

    Rebuilds the revenue forecast from pipeline, usage and bookings.

  • Books closed for the monthSnowflakeLookerGoogle Drive
    Delivers: Commentary for the CFO, with the numbers behind each driver
    Finance

    Variance Analysis

    Explains every material gap between budget and actuals, driver by driver.

  • Discount requested in SalesforceSalesforceSnowflakeSlack
    Delivers: Approve or counter, with the margin impact and precedent deals
    Finance

    Deal Desk Review

    Checks every non-standard deal against pricing policy and margin targets.

  • Contract signed in DocuSignDocuSignSalesforceGoogle Drive
    Delivers: The treatment for each contract, with the clauses behind it
    Finance

    Revenue Recognition Review

    Reads every signed contract for terms that change how its revenue is recognized.

  • Earnings release publishedWebGoogle DriveSlack
    Delivers: A one-page note on each company you cover within an hour of its release
    Finance

    Earnings Review

    Reads the release, the filing and the call transcript, and compares the results with your estimates.

  • Meeting on the calendarSalesforceGongGoogle Calendar
    Delivers: A one-page brief in the rep’s inbox an hour before every call
    Sales

    Account Research

    Builds a brief from the CRM, past calls, filings and news.

  • Deal moves to commitSalesforceGongSlack
    Delivers: The risks, missing stakeholders and next steps for the rep
    Sales

    Deal Review

    Reads every call and email in a late-stage deal and checks it against your sales process.

  • Daily 06:00SalesforceAmplitudeWeb
    Delivers: The accounts ready to expand, with the evidence for each
    Sales

    Expansion Signals

    Watches usage, hiring and news across your accounts for signs they’re ready to grow.

  • RFP added in SalesforceSalesforceGoogle DriveConfluenceSlack
    Delivers: A first draft of every answer, sourced from past wins, with gaps flagged
    Sales

    RFP Response

    Drafts RFP and security questionnaire answers from your approved content.

  • 90 days before renewalSalesforceAmplitudeGong
    Delivers: A renewal brief for each account: usage, open issues, sentiment and a plan
    Sales

    Renewal Risk Review

    Reads usage, account history and call notes before every renewal.

  • Mondays 08:00WebGongNotionSlack
    Delivers: A weekly brief on competitor moves and what they mean for open deals
    Strategy

    Competitive Intelligence

    Monitors competitor sites, pricing pages, job posts and mentions in your sales calls.

  • Data room sharedBoxGoogle DriveNotion
    Delivers: A diligence memo with every red flag sourced to the document and page
    Strategy

    Deal Diligence

    Reads the whole data room: financials, customer contracts, IP, employment and litigation.

  • Daily 07:00WebNotionSlack
    Delivers: New targets worth a look, with the case for each one
    Strategy

    Acquisition Screening

    Screens newly funded companies in your target markets against your criteria.

  • End of each quarterGongSalesforceNotion
    Delivers: A win-loss report with the calls behind every reason
    Strategy

    Win-Loss Analysis

    Reads the calls, notes and competitor mentions from every closed deal.

  • Two weeks before each board meetingSnowflakeSalesforceGoogle Drive
    Delivers: A first draft of the deck, every number tied to its source
    Strategy

    Board Deck Prep

    Pulls the quarter’s metrics, wins and risks into your board deck template.

  • API call from the procurement portalSalesforceDocuSignGoogle Drive
    Delivers: Approve, approve with conditions or reject, with evidence for every check
    Procurement

    Vendor Risk Review

    Runs sanctions, legal, security, financial and reference checks in parallel.

  • Daily 07:00WebGoogle DriveSlack
    Delivers: An alert with the evidence when a supplier becomes a risk
    Procurement

    Supplier Risk Watch

    Monitors key suppliers for financial distress, sanctions and outages.

  • Sourcing event closesGoogle DriveBoxSlack
    Delivers: A scored comparison and a recommended supplier, with the trade-offs
    Procurement

    Bid Comparison

    Compares supplier proposals line by line against your requirements and budget.

  • 60 days before a contract renewsDocuSignSnowflakeWeb
    Delivers: Your leverage, a target price and a walk-away point for every term
    Procurement

    Negotiation Prep

    Builds a negotiation brief from past spend, contract terms and market benchmarks.

  • Month-end closeSnowflakeLookerSlack
    Delivers: Savings ranked by value, with the spend behind each one
    Procurement

    Spend Analysis

    Reviews each month’s spend for duplicate vendors, price creep and missed discounts.

  • A scheduleAn alertAn eventAn API call
    And many more

    Any job that keeps coming back

    These are a few of the jobs teams run. Describe yours, what starts it and what good looks like, and an agent takes it from there.

Give it a job. Get back finished work.

Describe the job, what starts it and what a good result looks like. From then on Recursion runs it in the background: it plans the work, brings in as many specialist agents as it needs, runs them in parallel and delivers the result into your apps, with the evidence attached. Models, sandboxes, credentials, memory and grading are handled for you, so most teams use it out of the box.

You decide

Experiment Sweep
Job
Run every config in the spec and find the best one
Apps
LinearGitHubDatabricksSlack
Starts
Issue labeled sweep
Rubric
  • Every config runs to completion
  • Each run compared with the same baseline
  • The winner is rerun to confirm

Recursion runs

  • ModelsAny model, chosen per agent
  • SandboxesFor running code and using a browser
  • CredentialsKept in a vault, scoped to each tool
  • Multi-agentOne agent per task, and more for big jobs
  • MemoryNotes and skills from past runs
  • GradingEvery result checked against your rubric

You get

SlackSlack

48 configs ran in parallel overnight. Config 31 beats the baseline by 2.1 points, confirmed on a rerun.

  • Every config runs to completion
  • Each run compared with the same baseline
  • The winner is rerun to confirm
Graded 3 of 3

Every run lands on one fleet board.

Over time

Gets better, and cheaper, the more it works.

Every graded run leaves notes and skills the next run starts from. When a job runs often enough, Recursion turns its graded runs into training environments, the kind Labelbox builds for frontier labs, trains a specialist model on them and switches over once it beats the model you use today.

5–10×lower cost per task once a specialist takes over

Specialist takes over
GradeCost per task
MemoryBetter every runSpecialist modelBetter and cheaper

For frontier AI

The data and environments frontier models learn from

Over 90% of leading US AI labs post-train and evaluate on Labelbox.

Labelbox Research

Latest work from Labelbox Research

Our applied research team publishes the benchmarks and methods we use to train and evaluate frontier models.

Built for enterprise security

Each agent gets only the access its job needs, in a workspace of its own. The model never sees your passwords or keys, and every step is recorded.

  • SOC 2 Type II
  • GDPR
  • CCPA
  • NIST 800-171
  • PCI DSS
Security at Labelbox

Get started with Labelbox

For enterprises

Stop prompting. Start delegating.

Your teams already ask AI for answers. Recursion takes on whole jobs: it starts on its own, works across your apps and hands back finished, graded work. Bring the job that eats your experts’ week, and we’ll set it running.

For frontier AI

Build what your models learn next

Frontier AI labs already train on Labelbox environments and expert data. Tell us where your models fall short, and we’ll scope the environments, data and evals.