Prepare controls for an interpretability claim
Does a proposed attribution or activation-based conclusion survive its stated baseline, precision and intervention controls?
Make this plan your own ↓Read, edit and export without an account. This is preparation; no run or result is claimed.
What to compare
Define what observation would contradict the proposed explanation before inspecting outcomes. Distinguish correlation of activations from effects of an intervention, and keep numerical precision changes separate from scientific conclusions.
Useful outputs
- An exact model/checkpoint/tokenizer and a rights-appropriate fixed input set.
- A named attribution/intervention method, baseline and tensor/precision conventions.
- A frozen set of null, replacement and numerical controls tied to a narrowly stated claim.
Inputs and prerequisites
- Inspect current pinned library APIs and checkpoint/model/data terms independently.
- Qualify the tensor/GPU/dependency environment and quote its maximum budget before execution.
- The original numeric/synthetic examples illustrate control design; they do not validate these model-level methods.
What this work would not establish
- This is a planning packet with no model execution or interpretability finding.
- A visually plausible attribution map does not establish a causal mechanism.
Start with these sources.
TransformerLens
Transformer inspection tooling for activation capture and controlled interventions in mechanistic interpretability studies.
NNsight
Neural-network inspection and intervention interfaces for tracing internal activations and testing counterfactual computations.
SAE Lens
Sparse-autoencoder research tooling for examining reconstruction, sparsity and feature interventions on neural activations.
Captum
Model interpretability algorithms for checking attribution behavior against simple functions with known contributions.
PyTorch
Tensor computation and automatic differentiation framework for constructing controlled learning and gradient experiments.
Rationalization and subtractive cancellation
Measure loss of precision in sqrt(x+1)-sqrt(x) at x=10^16 and compare an algebraically rationalized expression with a high-precision reference.
Synthetic correlated measurements with a known covariance
Generate 256 paired synthetic measurements through a declared linear transformation of independent standard-normal draws, with known population covariance.
Make the question your own.
Edit freely. Export a copy before reloading or leaving. Your text stays in this tab until you choose to copy, download or save it privately.