01 Source
Source the knowledge. Vetted experts and enterprise data owners, matched to the capability you need.
- Credential-verified experts
- Enterprise data partners
- Matched to your capability spec
corpuslab builds training and evaluation data for frontier AI labs, created by vetted experts or licensed from the enterprises that own it, with provenance on every record.
01What we do
corpuslab connects the people and enterprises who hold scarce knowledge with the labs building frontier models, and turns that knowledge into data a model can learn from.
Demonstrations, reasoning traces and preference judgments written by vetted specialists, not crowds approximating them.
Proprietary text, audio, video, code and workflow data from enterprises, licensed under clear rights and usage terms.
Rubrics, verifiers, benchmarks and RL environments that measure real capability and reward the right behavior.
Every record carries its source, consent and license, so you always know what your model learned from.
02Platform
Source the knowledge. Vetted experts and enterprise data owners, matched to the capability you need.
License with clarity. Rights, consent and usage terms are set before a single record moves.
Verify every record. Expert review, automated checks and gold sets before anything ships.
Deliver to your stack. Versioned datasets with data cards and provenance, shipped to your bucket or API.
03Console
Traceable. Every record links to its source, consent and license.
Expert-led. Specialists write, review and score every task.
Private by default. Programs are isolated per customer.
Fast to first batch. Pilot data in days, not quarters.
04Who it's for
Built to spec. Every program starts from your capability gap.
For AI labs
Tell us where your model breaks. We scope the program, source experts and licensed data, and ship versioned datasets with provenance attached.
Owner-first licensing. You approve every buyer and every use.
For enterprises
We value your data, de-identify it and license it to frontier labs under terms you approve. You keep ownership. We handle buyers, delivery and payouts.
Expert-first work. Clear rates, flexible hours, credit for quality.
For experts
Physicians, lawyers, engineers, scientists and linguists write and review the data frontier models learn from. Rates are set by domain and shown before you start.
05Principles
No record enters a corpus without documented consent and clear rights.
Applies to: Labs · Owners · Experts
Licensing never transfers ownership. Buyers, terms and renewals stay your call.
Applies to: Owners
Rates are posted upfront and reflect the expertise the work demands.
Applies to: Experts
Every batch ships with its review scores, agreement metrics and known limits.
Applies to: Labs
We serve many labs and never share one customer's data or specs with another.
Applies to: Labs
Every record keeps its source, license and history for as long as it exists.
Applies to: Labs · Owners · Experts
From olympiad proofs to egocentric video.
Prove fit on a scoped first batch.
Pricing: Fixed scope
Dedicated experts and data for a capability target.
Pricing: Scoped per program
Measure capability before you train for it.
Pricing: Per benchmark or environment
Rights-cleared proprietary data from enterprise partners.
Pricing: Per license, term-based
Learn what your data is worth to frontier labs.
Pricing: No-cost assessment
License to several labs. Keep full ownership.
Pricing: Revenue share
One lab, premium terms, a defined window.
Pricing: Minimum guarantee + revenue share
Get paid for what you know.
Pricing: Hourly or per task
Questions, answered
Two places: work created by vetted experts under contributor agreements, and proprietary data licensed from the organizations that own it. Every record carries its source, consent basis and license terms. We don't sell scraped data of unclear origin.
We assess what's licensable, prepare it and offer it to labs under terms you approve. Each license pays a revenue share; exclusive deals add a guaranteed minimum. You see every buyer and every use, keep ownership, and decide on renewal at the end of each term.
Data is de-identified before delivery, using automated detection plus expert review for direct and indirect identifiers. Sensitive categories need the owner's explicit approval. Buyers receive only the approved, de-identified dataset, never raw source data.
A qualified expert writes each task, and a second expert reviews it. Each task is also checked against gold sets and automated validators. We track agreement, rubric pass rates and revision rates for each batch, and every delivery ships with a quality report.
Most programs begin with a scoped pilot. Sample data typically arrives within days and a first batch within weeks, depending on domain, volume and review depth.
JSONL, Parquet, CSV, WebDataset and Hugging Face-compatible datasets, or your own schema. We deliver to your S3, GCS or Azure bucket or through an API, with versioned releases, data cards and a provenance manifest.
Experts verify their identity and credentials, then pass domain assessments before joining a project. Their work is quality-scored on an ongoing basis. Pay is hourly or per task, set by domain, shown before you start, and paid on a fixed schedule.
Programs are isolated per customer. Access is least-privilege and logged, and data is encrypted in transit and at rest. Our controls are built for SOC 2 and GDPR-aligned workflows, and we sign DPAs and NDAs as standard.