A statistic, not a number
Agreement is reported with the method named, the item and rater counts stated, and a 95% interval from a seeded bootstrap — so the same inputs give you the same figure, every time.
For frontier and tier-2 AI research labs
PhD scientists, attorneys, clinicians and engineers doing the judgment work your evals depend on — measured by a team that doesn't run the work, and released with a signed evidence pack you can verify on your own machine.
Expert-only RLHF, SFT, evaluation and red-teaming programs — staffed, run and validated for the models you're training next.
Costed scope before you commit. NDA and security review first, if you prefer.
Proof, up front
Agreement is reported with the method named, the item and rater counts stated, and a 95% interval from a seeded bootstrap — so the same inputs give you the same figure, every time.
Released evidence packs are signed (ES256) against a publicly readable manifest, and ship with a standard-library Python script that re-computes them without us.
Deepen AI holds SOC 2 Type II, ISO 27001 and TISAX company-wide and co-authored ASAM OpenLABEL. DeepenSkill is being added to that program, not separately certified.
Expertise is a signal, not a volume game.
The problem
Models now clear the tasks a general contributor pool can evaluate. What's left needs someone who has actually done the work — the oncologist who can say why the staging is wrong, the litigator who spots the clause that voids the indemnity, the physicist who knows the derivation is unsound three lines before the error.
That expertise is hard to buy, and harder to check. Practising specialists aren't on labeling platforms; a résumé isn't a credential and a credential isn't competence. And when the vendor who recruited the expert is also the only party attesting to their work, the quality claim never rises above the résumé it started as.
Every vendor says their experts are good. The question is who measured, and whether you can run the measurement yourself.
What you get
You sign one contract. Experts work against your guidelines and gold reference data, and the evidence of their performance comes back with the work.
Specialist recruiting firms, BPOs and individual practitioners find the people — an aggregated network, not one pool.
A dedicated operations layer equips and runs each program — project management, workforce management and senior subject-matter oversight — so experts work to spec from day one.
A validation team separate from the team running the work scores it against gold reference data and reports how your experts are actually performing, back to you.
Never a roster to manage, never a supply chain to audit — and never a quality claim you have to take on trust.
Our network spans
How we prove it
Expert work is normally sold on attestation: the vendor says these people are excellent, and you find out in your evals eight weeks later. A vendor who only matches has no way around this — the only party who can vouch for their experts is the party that recruited them. DeepenSkill separates the team running the work from the team measuring it, and then hands you the means to check the measurement yourself.
Validation reports are authored by Deepen validators, and the platform rejects a validation report authored by the party that delivered the work. The independence statement is only attached to a report when every quality field on it was computed from tool metrics rather than typed by an operator — and when it wasn't, the pack says so.
Percent agreement, Cohen's κ, Krippendorff's α, Gwet's AC1 and AC2, weighted κ, gold-set accuracy, defect rate and rubric pass rate — each with the method named, the item and rater counts stated, and a 95% interval from a seeded bootstrap or a Wilson score interval for proportions. What each statistic measures, and when each one breaks →
The hypotheses, the sample, the statistics and the stopping rules are published and dated before a single judgment exists, so the result cannot be shaped after the fact. Read the pre-registration for our RewardBench 2 run →
A released evidence pack is signed with ES256 over a public-safe manifest — no names, no economics — and the manifest is readable from a public verify endpoint. There is no route that re-signs a pack, by construction.
verify-pack.py recomputes the pack hash and checks the signature against the published public key. verify-bundle.py goes further and recomputes every statistic from the raw judgment and defect records the bundle carries. Both are standard-library Python: nothing to install, nothing to trust but arithmetic you can read.
The audit log is hash-chained and anchored hourly into S3 Object Lock under a one-year governance retention; audit and evidence tables have UPDATE and DELETE revoked at the database-role level and again by trigger.
Credentials verified, a standardized domain test passed, then a practical task scored against gold reference data — enforced in the service layer so no stage can be skipped, with every advance written to a vetting record. The bar does not move depending on who introduced the person.
SOC 2 Type II · ISO 27001 · TISAX · GDPR-aligned.
Verifiable at security.deepen.aiFive onboarding items and two hard gates — six compliance attestations and confidentiality flow-down to every named individual among them — before one expert can be routed to your work.
The same bar for every partner — see it in fullThe honest caveat. Deepen AI's certifications are company-wide. DeepenSkill is being added to that existing program, not certified separately — there is no DeepenSkill-specific SOC 2 or TISAX report today, and we will not claim one until the scope extension is complete.We would rather lose a questionnaire than win it on a sentence that isn't true yet.
How it works
A costed, time-bounded estimate before you sign anything.
Through an aggregated network of recruiting firms, BPOs and practitioners — not one pool.
Credential verification, a standardized domain test, and a practical task scored against gold reference data. Stages cannot be skipped.
A dedicated ops pod — project management, workforce management and senior subject-matter oversight.
Against written guidelines, gold reference data, tiered review and formal arbitration.
Measured independently of the team running the work, released as a signed evidence pack, and verifiable by you.
Who it's for
Post-training and RLHF leads whose preference data has to come from people qualified to have the opinion. Evaluation and benchmark owners who need held-out sets that don't leak and don't drift. Safety and alignment teams running adversarial work in a real domain. Research-ops and data-procurement teams who have to defend a supplier choice to a security reviewer.
Frontier and tier-2 labs alike — the constant isn't scale, it's that a wrong judgment is expensive and a batch average won't tell you it happened.
Who it isn't for: high-volume generalist labeling where a large contributor pool is genuinely the right tool. We'll say so on the first call rather than sell you an expert bench you don't need. And running three expert vendors in rotation costs more than three invoices — it costs three vetting standards, three security reviews, and researcher hours spent re-screening people someone else already called screened.
Security & trust
Your task specs and clips live in encrypted S3 with public access blocked; every read and write goes through a short-lived presigned URL, never a public object URL. Services and the database run in private subnets inside one VPC, and every public hostname is TLS. Credentials and session secrets are generated straight into AWS Secrets Manager. Each service runs as a least-privilege database role with no schema rights, and the audit and evidence tables have update and delete revoked at the role level and again by trigger.
Partner organisation names and expert identities are stripped from lab-facing fields by a redaction function in code — experts and partners present to you under a Deepen-controlled identity, not their own. In production, the database is Multi-AZ with 14-day backups and point-in-time recovery, with CloudTrail, GuardDuty, ALB access logs and alarms on top.
Sub-processors are AWS for hosting, Google for sign-in only, AWS SES for transactional email, Stripe for invoicing, and vetted expert-sourcing partners as a category, whose named entities go to your DPO under NDA. Retention and deletion commitments are set in the DPA. Security incidents: skill@deepen.ai.
We send the full control list before a pilot, gaps included. The security overview we share names each control, cites the code or infrastructure that implements it, and states plainly what is live, what is planned and what does not exist yet.Deepen AI's certification detail is published at security.deepen.ai.
Read the full Trust Center — architecture, controls, sub-processors and the gaps →
Credibility
DeepenSkill is Deepen AI's expert-workforce platform — the same engineering, operations and compliance organization that has spent eight years delivering data infrastructure to customers who audit their suppliers: BMW, Aptiv, Bosch, Cadence and Daimler Trucks. Co-author of ASAM OpenLABEL.
FAQ
The team that measures the work does not run it. Validation reports are authored by Deepen validators, and the platform rejects a validation report authored by the party that delivered the work. That separation is enforced in the product, not promised in a slide.
Percent agreement, Cohen's κ, Krippendorff's α, Gwet's AC1 and AC2, weighted κ, gold-set accuracy, defect rate and rubric pass rate — each with the method named, the item and rater counts stated, and a 95% interval (seeded bootstrap, or a Wilson score interval for proportions). A point estimate on its own is not a result.
Yes. A released evidence pack carries an ES256 signature and a publicly readable manifest. Two standard-library Python scripts recompute the pack hash, check the signature against the published public key, and recompute every statistic from the underlying judgment records on your own machine — with no dependency on our dashboard. The public sample on this site demonstrates the statistics half (a verification bundle you can recompute); pack signing happens at release on the platform. The public sample on this site demonstrates the statistics half (a verification bundle you can recompute); pack signing happens at release on the platform.
One bar for everyone, regardless of who introduced them: credentials verified, a standardized domain test passed, then a practical task scored against gold reference data before anyone becomes active. The stages are enforced in the service layer and cannot be skipped, and each advance is written to a vetting record.
No, unless you explicitly agree to an arrangement where it does. Sourcing partners handle expert records — identity, credentials and payment detail — not your task material. In the other direction, partner organisation names and expert identities are stripped from lab-facing fields by a redaction function in code, so experts and partners present to you under a Deepen-controlled identity.
Deepen AI, the parent company, holds SOC 2 Type II, ISO 27001 and TISAX company-wide, with a GDPR-aligned data-protection posture. DeepenSkill is being added to that existing program, not certified separately — there is no DeepenSkill-specific report today, and we will not represent one until that scope extension is complete.
Definitions for the terms above: RLHF, gold standard, provenance and model evaluation — or the full glossary.
Get started
Bring the task, the bar you need to hit, and the reason your current source isn't hitting it. We'll come back with a costed scope, a quality plan and the metric it has to beat — and we'll show you the evidence format you'd be receiving before you commit to anything. If we're the wrong supplier for that mandate, we'll say so on the first call.
Refinery is where you license data. DeepenSkill is where you engage people.