A scientist working at a whiteboard

For frontier and tier-2 AI research labs

Expert talent, with the proof attached.

PhD scientists, attorneys, clinicians and engineers doing the judgment work your evals depend on — measured by a team that doesn't run the work, and released with a signed evidence pack you can verify on your own machine.

Expert-only RLHF, SFT, evaluation and red-teaming programs — staffed, run and validated for the models you're training next.

Costed scope before you commit. NDA and security review first, if you prefer.

Proof, up front

Three things no talent vendor hands you.

A statistic, not a number

Agreement is reported with the method named, the item and rater counts stated, and a 95% interval from a seeded bootstrap — so the same inputs give you the same figure, every time.

A signature you can check

Released evidence packs are signed (ES256) against a publicly readable manifest, and ship with a standard-library Python script that re-computes them without us.

A company procurement already knows

Deepen AI holds SOC 2 Type II, ISO 27001 and TISAX company-wide and co-authored ASAM OpenLABEL. DeepenSkill is being added to that program, not separately certified.

Macro view of a firing neuron synapse
Expertise is a signal, not a volume game.

The problem

The frontier moved past what a crowd can grade.

Models now clear the tasks a general contributor pool can evaluate. What's left needs someone who has actually done the work — the oncologist who can say why the staging is wrong, the litigator who spots the clause that voids the indemnity, the physicist who knows the derivation is unsound three lines before the error.

That expertise is hard to buy, and harder to check. Practising specialists aren't on labeling platforms; a résumé isn't a credential and a credential isn't competence. And when the vendor who recruited the expert is also the only party attesting to their work, the quality claim never rises above the résumé it started as.

Every vendor says their experts are good. The question is who measured, and whether you can run the measurement yourself.

What you get

Experts working to your spec — and the record of how they did.

You sign one contract. Experts work against your guidelines and gold reference data, and the evidence of their performance comes back with the work.

The work: preference and RLHF judgments from genuine domain experts, reward-model and rubric authoring, expert evaluation and held-out benchmark sets, adversarial red-teaming in domain, long-form reasoning traces and demonstrations, human-in-the-loop review and adjudication of the cases your pipeline can't settle, and multimodal annotation and labeling where the judgment matters more than the throughput.

Source

Specialist recruiting firms, BPOs and individual practitioners find the people — an aggregated network, not one pool.

Equip

A dedicated operations layer equips and runs each program — project management, workforce management and senior subject-matter oversight — so experts work to spec from day one.

Validate

A validation team separate from the team running the work scores it against gold reference data and reports how your experts are actually performing, back to you.

Never a roster to manage, never a supply chain to audit — and never a quality claim you have to take on trust.

Editorial still life of artifacts from specialist domains
The disciplines behind the work are not interchangeable.

Our network spans

  • Medicine & life sciences
  • Law & regulatory
  • Physics & mathematics
  • Software & security
  • Advanced engineering
  • Finance & actuarial
Molecular structure floating in deep space
Signal · Evidence · Provenance

How we prove it

Don't trust our dashboard. Run our script.

Expert work is normally sold on attestation: the vendor says these people are excellent, and you find out in your evals eight weeks later. A vendor who only matches has no way around this — the only party who can vouch for their experts is the party that recruited them. DeepenSkill separates the team running the work from the team measuring it, and then hands you the means to check the measurement yourself.

  • The measurement is not made by the producer

    Validation reports are authored by Deepen validators, and the platform rejects a validation report authored by the party that delivered the work. The independence statement is only attached to a report when every quality field on it was computed from tool metrics rather than typed by an operator — and when it wasn't, the pack says so.

  • Agreement with a method, an interval and a seed

    Percent agreement, Cohen's κ, Krippendorff's α, Gwet's AC1 and AC2, weighted κ, gold-set accuracy, defect rate and rubric pass rate — each with the method named, the item and rater counts stated, and a 95% interval from a seeded bootstrap or a Wilson score interval for proportions. What each statistic measures, and when each one breaks →

  • We pre-register our benchmark runs

    The hypotheses, the sample, the statistics and the stopping rules are published and dated before a single judgment exists, so the result cannot be shaped after the fact. Read the pre-registration for our RewardBench 2 run →

  • Signed packs, and the bytes you download are the bytes we signed

    A released evidence pack is signed with ES256 over a public-safe manifest — no names, no economics — and the manifest is readable from a public verify endpoint. There is no route that re-signs a pack, by construction.

  • Buyer-runnable verification

    verify-pack.py recomputes the pack hash and checks the signature against the published public key. verify-bundle.py goes further and recomputes every statistic from the raw judgment and defect records the bundle carries. Both are standard-library Python: nothing to install, nothing to trust but arithmetic you can read.

  • An audit log we can't rewrite either

    The audit log is hash-chained and anchored hourly into S3 Object Lock under a one-year governance retention; audit and evidence tables have UPDATE and DELETE revoked at the database-role level and again by trigger.

  • One vetting bar, applied to everyone

    Credentials verified, a standardized domain test passed, then a practical task scored against gold reference data — enforced in the service layer so no stage can be skipped, with every advance written to a vetting record. The bar does not move depending on who introduced the person.

Deepen AI holds

SOC 2 Type II · ISO 27001 · TISAX · GDPR-aligned.

Verifiable at security.deepen.ai

Sourcing partners clear

Five onboarding items and two hard gates — six compliance attestations and confidentiality flow-down to every named individual among them — before one expert can be routed to your work.

The same bar for every partner — see it in full

The honest caveat. Deepen AI's certifications are company-wide. DeepenSkill is being added to that existing program, not certified separately — there is no DeepenSkill-specific SOC 2 or TISAX report today, and we will not claim one until the scope extension is complete.We would rather lose a questionnaire than win it on a sentence that isn't true yet.

How it works

Six steps. You see the evidence at every one.

  1. Scope

    A costed, time-bounded estimate before you sign anything.

  2. Source

    Through an aggregated network of recruiting firms, BPOs and practitioners — not one pool.

  3. Vet

    Credential verification, a standardized domain test, and a practical task scored against gold reference data. Stages cannot be skipped.

  4. Equip

    A dedicated ops pod — project management, workforce management and senior subject-matter oversight.

  5. Produce

    Against written guidelines, gold reference data, tiered review and formal arbitration.

  6. Prove it

    Measured independently of the team running the work, released as a signed evidence pack, and verifiable by you.

Who it's for

Teams whose next result depends on a judgment a crowd can't make.

Post-training and RLHF leads whose preference data has to come from people qualified to have the opinion. Evaluation and benchmark owners who need held-out sets that don't leak and don't drift. Safety and alignment teams running adversarial work in a real domain. Research-ops and data-procurement teams who have to defend a supplier choice to a security reviewer.

Frontier and tier-2 labs alike — the constant isn't scale, it's that a wrong judgment is expensive and a batch average won't tell you it happened.

Who it isn't for: high-volume generalist labeling where a large contributor pool is genuinely the right tool. We'll say so on the first call rather than sell you an expert bench you don't need. And running three expert vendors in rotation costs more than three invoices — it costs three vetting standards, three security reviews, and researcher hours spent re-screening people someone else already called screened.

What matters DeepenSkill Typical vendor
Quality evidence Item-level measurement Résumé or badge
Who measures Separate from the producer Internal or unclear
Buyer can re-verify Signed pack + script Vendor dashboard
Domain depth Practising experts Generalist coverage

Security & trust

Built for a buyer whose security team asks first.

Your task specs and clips live in encrypted S3 with public access blocked; every read and write goes through a short-lived presigned URL, never a public object URL. Services and the database run in private subnets inside one VPC, and every public hostname is TLS. Credentials and session secrets are generated straight into AWS Secrets Manager. Each service runs as a least-privilege database role with no schema rights, and the audit and evidence tables have update and delete revoked at the role level and again by trigger.

Partner organisation names and expert identities are stripped from lab-facing fields by a redaction function in code — experts and partners present to you under a Deepen-controlled identity, not their own. In production, the database is Multi-AZ with 14-day backups and point-in-time recovery, with CloudTrail, GuardDuty, ALB access logs and alarms on top.

Sub-processors are AWS for hosting, Google for sign-in only, AWS SES for transactional email, Stripe for invoicing, and vetted expert-sourcing partners as a category, whose named entities go to your DPO under NDA. Retention and deletion commitments are set in the DPA. Security incidents: skill@deepen.ai.

We send the full control list before a pilot, gaps included. The security overview we share names each control, cites the code or infrastructure that implements it, and states plainly what is live, what is planned and what does not exist yet.Deepen AI's certification detail is published at security.deepen.ai.

Read the full Trust Center — architecture, controls, sub-processors and the gaps →

Credibility

New platform. Not a new company.

DeepenSkill is Deepen AI's expert-workforce platform — the same engineering, operations and compliance organization that has spent eight years delivering data infrastructure to customers who audit their suppliers: BMW, Aptiv, Bosch, Cadence and Daimler Trucks. Co-author of ASAM OpenLABEL.

FAQ

The six questions buyers actually ask.

  • How is your quality claim different from a vendor vouching for its own people?

    The team that measures the work does not run it. Validation reports are authored by Deepen validators, and the platform rejects a validation report authored by the party that delivered the work. That separation is enforced in the product, not promised in a slide.

  • What agreement statistics do you report, and how?

    Percent agreement, Cohen's κ, Krippendorff's α, Gwet's AC1 and AC2, weighted κ, gold-set accuracy, defect rate and rubric pass rate — each with the method named, the item and rater counts stated, and a 95% interval (seeded bootstrap, or a Wilson score interval for proportions). A point estimate on its own is not a result.

  • Can we verify the numbers ourselves?

    Yes. A released evidence pack carries an ES256 signature and a publicly readable manifest. Two standard-library Python scripts recompute the pack hash, check the signature against the published public key, and recompute every statistic from the underlying judgment records on your own machine — with no dependency on our dashboard. The public sample on this site demonstrates the statistics half (a verification bundle you can recompute); pack signing happens at release on the platform. The public sample on this site demonstrates the statistics half (a verification bundle you can recompute); pack signing happens at release on the platform.

  • Who vets the experts, and to what bar?

    One bar for everyone, regardless of who introduced them: credentials verified, a standardized domain test passed, then a practical task scored against gold reference data before anyone becomes active. The stages are enforced in the service layer and cannot be skipped, and each advance is written to a vetting record.

  • Does our task data reach your sourcing partners?

    No, unless you explicitly agree to an arrangement where it does. Sourcing partners handle expert records — identity, credentials and payment detail — not your task material. In the other direction, partner organisation names and expert identities are stripped from lab-facing fields by a redaction function in code, so experts and partners present to you under a Deepen-controlled identity.

  • What certifications does DeepenSkill hold?

    Deepen AI, the parent company, holds SOC 2 Type II, ISO 27001 and TISAX company-wide, with a GDPR-aligned data-protection posture. DeepenSkill is being added to that existing program, not certified separately — there is no DeepenSkill-specific report today, and we will not represent one until that scope extension is complete.

Definitions for the terms above: RLHF, gold standard, provenance and model evaluation — or the full glossary.

Get started

Ask us for the evidence before you ask us for a quote.

Bring the task, the bar you need to hit, and the reason your current source isn't hitting it. We'll come back with a costed scope, a quality plan and the metric it has to beat — and we'll show you the evidence format you'd be receiving before you commit to anything. If we're the wrong supplier for that mandate, we'll say so on the first call.

Refinery is where you license data. DeepenSkill is where you engage people.