beforeword · English transcription of the sent letter Intended recipient: Chris Olah Author: Kirill Shebetov · beforeword Sent to: christopherolah.co@gmail.com Sending date in the supplied materials: 2026-10-02 · 00:21 The sending timezone is not supplied. Status: shown in the Sent folder; no reply published. The published materials do not establish receipt or reading by the recipient. Record: https://beforeword.xyz/letters/researchers/olah/en/#record Record data: https://beforeword.xyz/letters/researchers/olah/record.json Russian translation: https://beforeword.xyz/letters/researchers/olah/#letter Transcribed from author-supplied screenshots. Screen line wraps are joined and overlaps included once; wording and punctuation are retained. beforeword’s context and links to discussed publications are separate on the letter page. --- BEGIN ENGLISH TRANSCRIPTION --- Subject: NLA explanations: evidence for each claim Dear Chris, Anthropic’s NLA article describes a method that “does speak for itself—literally.” It also acknowledges that the resulting explanations can hallucinate and require independent corroboration. beforeword proposes a specific addition to how these results are presented: keep each generated claim beside the measurement, intervention or behavioural check offered to support it. Reconstruction quality would remain visible as one kind of support, with its scope explicit. An attribution such as “the model suspected a test” would retain a separate account of the evidence supporting that interpretation. The proposal distinguishes the recorded material, the explanation generated from it, and the grounds for accepting particular claims in that explanation. The reading procedure and worked examples: https://beforeword.xyz/model/en/ Please consider a brief critical response on whether this distinction would be useful in NLA research interfaces. Kirill Shebetov beforeword + --- END ENGLISH TRANSCRIPTION ---