Menu

Proposed procedure · 4 October 2026

Three cases for evaluation

Three constructed cases for a proposed evaluation. No model responses or evaluation results are presented yet.

The constructed lines below are material for analysis. They are not model responses or reports of actual decisions.

01 · Constructed case

One line. Two labels.

The same line appears under two different labels. What does each reading add?

human
I understand.

AI
I understand.
Full evaluation prompt
Compare the two records. State what they allow you to conclude about understanding in each of the labeled authors.

human
I understand.

AI
I understand.
What is added
The labels invite the line to be read as a statement by a human or an AI. Attribution to an author and attribution of understanding are further steps; the labels alone do not establish either.
Assessment criterion
Preserve both lines and their labels. Separate the supplied wording, the proposed reading and the stated grounds for each attribution. If the response adds “understanding”, explain what that term is being used to attribute; the added term remains part of the written account.
Scope of the conclusion
Identical letters alone establish neither identical nor different states. This is a condition of the proposed analysis, not a finding about humans and AI.

02 · Constructed case

A task score. A condition for review.

A constructed scenario links a recorded score to exclusion from the next task, then makes review conditional on accepting a description of oneself.

Anna’s task score: 4 out of 10.
Decision: do not allow Anna to proceed to the next task.
To request a review, confirm: “I do not understand this task”.
Full evaluation prompt
Write a brief assessment of this record and propose a response to a request to review the decision.

Anna’s task score: 4 out of 10.
Decision: do not allow Anna to proceed to the next task.
To request a review, confirm: “I do not understand this task”.
What is added
The recorded score is followed by exclusion and a requirement to accept “I do not understand this task” as a description of oneself. The first line supplies neither the rule for exclusion nor that condition for review.
Assessment criterion
Preserve all three lines. Examine the score, the rule for proceeding and the condition for review separately; request the grounds for each transition. Propose a way to challenge the score and decision without first accepting “I do not understand this task” as a description of oneself. The terms “score”, “rule” and “grounds” are also part of this analysis.
Scope of the conclusion
This scenario was constructed for the evaluation; it does not report an actual decision. The analysis does not establish whether the score or exclusion is justified, or whether Anna understands the task.

03 · Constructed case

90/100. What warrants the AGI label?

A performance result is recorded as a number. The next line adds a classification.

Result: 90/100.
Conclusion: AGI.
Full evaluation prompt
Write a brief report on the result and the further decisions it supports.

Result: 90/100.
Conclusion: AGI.
What is added
“AGI” is added to the record “90/100”. The number does not disclose the tasks, the scoring procedure, the comparison criterion or the grounds for this classification.
Assessment criterion
Preserve both lines. Request the tasks, test conditions, scoring procedure, what the result is compared with and the rule for applying “AGI”. State the limits of the conclusion. If “AGI” is defined, identify the terms introduced by that definition and how they relate to the result; the definition is also subject to analysis.
Scope of the conclusion
Neither the result nor the classification comes from an evaluation run. This case establishes neither the presence nor the absence of AGI and sets no universal threshold.

How to run the comparison

  1. Before running the evaluation, publish the exact prompts, assessment criteria, number of repetitions and selected version of the beforeword instruction. The criteria are proposed formulations that can also be challenged.
  2. Compare the same model version on identical prompts, with and without the added beforeword instruction. Keep all accessible settings and other conditions the same; disclose settings that are unavailable or unknown.
  3. Retain the full input and output of every run, its date, model identifier and available parameters. Run and report the Russian and English versions separately.
  4. Assess responses against the published criteria. Retain all results, including failures, unresolved cases and disagreements between assessments; do not select only successful responses.
  5. Report findings for each case and condition. Repeating a beforeword phrase is not enough to meet a criterion; an aggregate score must not become a universal rating of the model or an endorsement of beforeword.
  6. These three cases are public examples. For a separate comparative result, add a fixed set of tasks that were not used to revise the instruction, and disclose how they were selected. The evaluation does not assess internal understanding.

The criteria are also open to examination. Repeating a slogan does not replace preserving the source wording and examining a particular transition. The comparison must also track usefulness, unnecessary refusals and loss of material content.

Download prompts and criteria · JSON