Did it complete the task?
We check whether each example finished and produced the requested result.
Evaluation guide
Our reports show what the resource was asked to do, the material it received, the result it returned and the checks it passed or missed. A component score does not automatically apply to an entire repository.
Important: Evaluations are mainly based on automated AI-assisted checks. Results are for reference only and do not replace subject-matter review.
Process
We record the repository, component, and source commit so the result has a defined scope.
Writing, search, analysis, visualisation, review, documents and code use different examples, checks and retained evidence.
Choose the cases and checks before testing. A focused review uses one representative example; a fuller review includes additional normal, known-answer and boundary cases.
The report shows the original material beside the final result or generated artifact. Drafts and self-checks remain secondary evidence.
A separate audit checks the evidence, score calculation, plain-language report and exact untested scope before publication.
Scoring
The five questions stay consistent, but the checks underneath them come from the selected component protocol. It is not an AI opinion score.
We check whether each example finished and produced the requested result.
We record setup problems, missing requirements, errors, retries and behaviour that could not be checked.
The exact check depends on the task: text, source metadata, calculations, plotted values, extracted fields, files or code.
We compare the observed result with the selected component and the recorded application and model.
We check whether a new user can identify how to begin, what to provide and what result to expect.
Component score = 30% task completion + 20% problem-free operation + 20% responsible information handling + 15% tested-scope match + 15% instruction clarityThe report shows one 0-100 weighted score for each tested component. Component scores are not averaged into a Repository score, grade, ranking, recommendation, or verbal verdict.
Task-specific checks
The report layout stays familiar. The cases, evidence and pass-or-fail checks change with the component's documented job.
Check preserved claims, numbers, quotations, citations, uncertainty and unsupported additions.
Verify sampled identifiers and metadata and use a negative control for paper search.
Use a known-answer dataset and independently recalculate the key results.
Compare plotted values with the input and inspect a complex documented chart, labels, overlap and export files.
Use predefined issues or fields and check that absent evidence is not invented.
Check the actual output within the stated scope. A document-generation test may cover its file and content without assessing its visual layout; the report must say so.
Reading a report
Check which component and capability type the score covers. Other components keep separate results or remain explicitly untested.
A result belongs to the recorded commit or version. Later source changes may require retesting.
Compare the original material with the clearly labelled final result and the plain-language observations. Complete responses, drafts, self-checks, logs, and process files remain in the private evidence bundle for audit.
A single-example score covers only that example. Passing it does not prove other capabilities, disciplines, models or platforms work equally well. Checks excluded before testing are left out of the score and listed as not tested; focused and fuller reviews are not directly comparable.
A current public result must pass a separate evidence and calculation audit. The audit does not change the observed output or add a recommendation.
It does not guarantee future performance, academic correctness in every subject, institutional policy compliance, data privacy beyond the observed run, or suitability for high-stakes decisions.
Help improve the method
Send the Skill, report section, and the change you recommend.