Menu

Published letter · AI and model evaluation

Chris Olah

What supports each claim in an explanation of a model’s behavior?

Sent · no reply published

Context by beforeword · separate from the sent text

Why this letter was written

What prompted the letter

The letter concerns NLA, a method used to produce explanations of a model’s workings. It responds to Anthropic’s description of the method as speaking for itself, alongside the article’s acknowledgment that the explanations need independent corroboration.

Why this matters to readers

Readers should be able to see what supports each claim: a measurement, an intervention in the model, or a behavioral check. This allows a claim such as “the model suspected a test” to be examined separately from the material offered in support.

What the letter asks

The letter proposes showing each generated claim alongside that material and asks whether this would help researchers use the results.

Publications and the passages discussed

Sent letter · English transcription

The letter

Transcribed from the supplied screenshots. Screen line breaks are joined into paragraphs; wording and punctuation are retained.

Dear Chris,

Anthropic’s NLA article describes a method that “does speak for itself—literally.” It also acknowledges that the resulting explanations can hallucinate and require independent corroboration.

beforeword proposes a specific addition to how these results are presented: keep each generated claim beside the measurement, intervention or behavioural check offered to support it.

Reconstruction quality would remain visible as one kind of support, with its scope explicit. An attribution such as “the model suspected a test” would retain a separate account of the evidence supporting that interpretation.

The proposal distinguishes the recorded material, the explanation generated from it, and the grounds for accepting particular claims in that explanation.

The reading procedure and worked examples:
https://beforeword.xyz/model/en/

Please consider a brief critical response on whether this distinction would be useful in NLA research interfaces.

Kirill Shebetov
beforeword

+

Publications discussed in the letter

These links are provided for reading in context. They are separate from the sent letter; beforeword’s commentary is not a reply from the recipient.

  1. Natural Language Autoencoders: Turning Claude’s thoughts into text

    7 May 2026. Introduction; What is a natural language autoencoder?; The future of NLAs.

    The article contains the discussed wording, the reconstruction-based evaluation method and limitations on the explanations.

Sending record and files

Letter author
Kirill Shebetov · beforeword
Sending date in the record
2 October 2026 · 00:21
Sent email subject
NLA explanations: evidence for each claim

The author supplied screenshots from the Sent folder. They show the message and the available sending details. These materials do not record the recipient reading it.

View screenshots · 2
  1. Letter to Chris Olah. Screenshot 1.

    Screenshot 1 · Open full size

  2. Letter to Chris Olah. Screenshot 2.

    Screenshot 2 · Open full size

Each TXT file begins with publication details, followed by the separately marked letter text. These are screenshot transcriptions and translations, not original email files. The purpose explanation and opening question are editorial summaries by beforeword.

Correspondence record

  1. Sending shown in the author’s materials

    2 October 2026 · 00:21

    Receipt by the intended recipient and a substantive reply are not separately recorded in the published materials.

  2. Added to the public collection

    The English transcription, Russian translation and screenshots are presented separately. No replies have been published yet.

The absence of a published reply does not establish agreement, rejection or an inability to object.