Evaluation guide

How AI Labour Department evaluations work

Our reports show what the resource was asked to do, the material it received, the result it returned and the checks it passed or missed. A component score does not automatically apply to an entire repository.

Important: Evaluations are mainly based on automated AI-assisted checks. Results are for reference only and do not replace subject-matter review.

Process

From source snapshot to report

  1. 01

    Fix the evaluation target

    We record the repository, component, and source commit so the result has a defined scope.

  2. 02

    Choose the test protocol

    Writing, search, analysis, visualisation, review, documents and code use different examples, checks and retained evidence.

  3. 03

    Fix the cases before running

    Choose the cases and checks before testing. A focused review uses one representative example; a fuller review includes additional normal, known-answer and boundary cases.

  4. 04

    Run and compare

    The report shows the original material beside the final result or generated artifact. Drafts and self-checks remain secondary evidence.

  5. 05

    Review the evaluation independently

    A separate audit checks the evidence, score calculation, plain-language report and exact untested scope before publication.

Scoring

Five evidence dimensions

The five questions stay consistent, but the checks underneath them come from the selected component protocol. It is not an AI opinion score.

Did it complete the task?

We check whether each example finished and produced the requested result.

Did it run without problems?

We record setup problems, missing requirements, errors, retries and behaviour that could not be checked.

Did it handle the supplied information responsibly?

The exact check depends on the task: text, source metadata, calculations, plotted values, extracted fields, files or code.

Did it do the documented task in the tested setup?

We compare the observed result with the selected component and the recorded application and model.

Were the instructions easy to follow?

We check whether a new user can identify how to begin, what to provide and what result to expect.

Component score = 30% task completion + 20% problem-free operation + 20% responsible information handling + 15% tested-scope match + 15% instruction clarity

The report shows one 0-100 weighted score for each tested component. Component scores are not averaged into a Repository score, grade, ranking, recommendation, or verbal verdict.

Task-specific checks

Different functions are tested differently

The report layout stays familiar. The cases, evidence and pass-or-fail checks change with the component's documented job.

Academic writing

Check preserved claims, numbers, quotations, citations, uncertainty and unsupported additions.

Literature and citations

Verify sampled identifiers and metadata and use a negative control for paper search.

Data analysis

Use a known-answer dataset and independently recalculate the key results.

Data visualisation

Compare plotted values with the input and inspect a complex documented chart, labels, overlap and export files.

Review and extraction

Use predefined issues or fields and check that absent evidence is not invented.

Documents, code and workflows

Check the actual output within the stated scope. A document-generation test may cover its file and content without assessing its visual layout; the report must say so.

Reading a report

What to check before relying on a result

Evaluation target

Check which component and capability type the score covers. Other components keep separate results or remain explicitly untested.

Source snapshot

A result belongs to the recorded commit or version. Later source changes may require retesting.

Input and output examples

Compare the original material with the clearly labelled final result and the plain-language observations. Complete responses, drafts, self-checks, logs, and process files remain in the private evidence bundle for audit.

Scope limits

A single-example score covers only that example. Passing it does not prove other capabilities, disciplines, models or platforms work equally well. Checks excluded before testing are left out of the score and listed as not tested; focused and fuller reviews are not directly comparable.

Independent review

A current public result must pass a separate evidence and calculation audit. The audit does not change the observed output or add a recommendation.

What an evaluation does not certify

It does not guarantee future performance, academic correctness in every subject, institutional policy compliance, data privacy beyond the observed run, or suitability for high-stakes decisions.

View the reference evaluation

Support a Skill

Send rockets to a Skill

One rocket costs $2 and adds 2 points.

Total1 rocket · $2 · 2 points

Rocket totals and aggregate revenue are public. Creators receive 70% of distributable net profit after payment fees, refunds, chargebacks, applicable taxes and payout costs.

You will review the total on Stripe before paying. There is no overall rocket limit; Stripe applies a technical limit to each checkout.