Interpretability as a First-Class Layer
neuros-mechint brings interpretability into the same platform as the models, so understanding a
representation is a built-in capability rather than a separate notebook pile.
package map
neuros-mechint — mechanistic interpretability + biophysics
Interpretability inside the platform: circuit discovery, sparse/concept features, representation alignment, dynamics, and biophysical grounding back to biology.
Circuits
ACDC + path patching
automated circuit discovery
Motif / feature viz / DUNL
structure and visualization
Counterfactuals + attribution
causal edits
Features & alignment
Sparse / concept SAEs
feature decomposition
RSA / CCA / PLS
representation alignment
Cross-species + temporal
aligned across brains and time
Grounding
Dynamics + bifurcation
state-space analysis
Biophysical models
ion channels, synapses, Dale, metabolic
What It Comprises
- Circuits — ACDC circuit discovery, path patching, DUNL, motif detection, feature
visualization, latent-RNN, and a circuit comparator; plus
attributionandcounterfactuals. - Concepts / features — sparse autoencoders and concept SAEs for feature decomposition.
- Representation alignment — RSA, CCA, PLS, temporal and cross-species alignment with validation and metrics.
- Dynamics — state-space dynamics analysis and bifurcation tooling.
- Biophysical grounding — compartmental and spiking-net models, ion channels, synaptic models, Dale's law, and metabolic constraints.
How It Differs
Where the other packages build models, neuros-mechint explains them — and uniquely reaches
back to biophysics, connecting learned artificial representations to real neural mechanisms. It is
the reflective half of the platform.
Part of the neurOS-v1 platform.