See what a model is doing below token level.
ModelCypher is a measurement and observability workbench for open-source model builders. It gives humans and frontier AI a clear way to inspect geometry, entropy, curvature, chain structure, and adapter-induced changes through workflow-first CLI surfaces instead of ad hoc activation scripts.
Current evidence state (reviewed 2026-07-11): mc analyze is the clearest public
entrypoint for prompt capture, prompt-family studies, and checkpoint or
adapter comparison. mc train run remains shipped and geometry-derived, but
the repo has not yet closed the promotable same-model same-data same-eval
benchmark needed to claim "better than standard practice." The retained 350M
R2 failure is closed as a data-format mechanism; it is not evidence of a
general inference-geometry collapse. See
RESEARCH-ROADMAP.md.
A forward pass is a deterministic geometric map, and mc analyze is the
shipped workbench for measuring that map below token level. The downstream
training program derives all 15 traditional controls from model geometry,
precision, or measured data where the current derivation is wired; the runtime
status table below names the places where the shipped default still uses a
calibrated AdamW value or an unwired formula. See
AGENTS.md for the full derivation philosophy.
poetry run mc analyze capture --model /path/to/model --prompt "Explain geodesics."
poetry run mc analyze family --model /path/to/model --manifest data/probes/prompt_family_minimal_pairs.json
poetry run mc analyze compare --left-model /path/to/base --right-model /path/to/base --right-adapter /path/to/adapter --manifest data/probes/prompt_family_minimal_pairs.json
poetry run mc analyze report --bundle /path/to/bundle
poetry run mc analyze report --bundle results/measurement_atlas
poetry run mc analyze report --bundle results/measurement_atlas/<run_id>
poetry run mc analyze report --bundle results/pipeline_validation
poetry run python scripts/run_measurement_atlas.py --model /path/to/model --manifest data/probes/measurement_atlas_casing.json --manifest data/probes/measurement_atlas_profanity_tone.json --manifest data/probes/measurement_atlas_grounded_hallucination.json --output-root results/measurement_atlasThese commands emit an observation bundle under
results/analysis/<timestamp-slug>/ by default:
manifest.jsonsummary.jsonREPORT.mdvariants.jsonllayer_metrics.jsonlcomparisons.jsonl
The prompt-family interface is explicit in phase 1. Each row includes:
case_id, variant_id, text, optional tags, and optional
comparison_to.
For research-only generation tracing, the measurement atlas runner writes a
family artifact under results/measurement_atlas/<run_id>/ with:
run_manifest.jsonsummary.jsonREPORT.mdledger.tsvvariants.jsonlsequence_metrics.jsonlstep_metrics.jsonlspace_step_metrics.jsonlcomparisons.jsonlonset_events.jsonl
The retained replay-alignment closure for the shipped 350M atlas pack has a
tracked report copy at
docs/research/reports/measurement_atlas/REPORT.md.
Current observed atlas surfaces are replay={hidden, embedding} and
live={hidden}; run_manifest.json now records requested vs observed
surfaces separately so the bundle does not overclaim unsupported replay space
coverage. mc analyze report --bundle ... can now read both the standard
mc analyze bundles, the retained atlas family root, individual atlas artifact
directories, and retained pipeline-validation families. Atlas generation itself
remains research-only in scripts/run_measurement_atlas.py.
poetry run mc train run --model /path/to/model --data /path/to/dataset --output /path/to/adapterThe shipped path derives the adapter surface, rank, batch sizing, scale budget, and stopping certificate from measured geometry. Its default optimizer is still the calibrated AdamW/cosine path called out in the status table; MASS remains available on research modes until the closure benchmark earns promotion.
Need extra instrumentation? Use flags on the same command path, such as
--benchmark, --topo-monitor, --dim-monitor, or --entropy-reg.
| # | Training control | Runtime status | Current code truth | Code source |
|---|---|---|---|---|
| 1 | Learning rate | derived+research-mode-only | default: calibrated AdamW 2e-4 cosine; MASS on research modes | _mlx_training_adapter_train_mixin.py + mass_step_size.py |
| 2 | Adam epsilon | removed | no Adam epsilon derivation is claimed; compute_geometric_epsilon is ScaledGD regularization in weight-spectrum units |
geometric_optimizer.py |
| 3 | Momentum | derived+research-mode-only | default: AdamW betas 0.9/0.999; Fisher/MASS moments only in research modes | _mlx_training_adapter_train_mixin.py + diagonal_fisher_preconditioner.py |
| 4 | Weight decay | formula-exists-unwired | condition-ratio formula exists; shipped mc train run default passes weight_decay=0.0 |
geometric_optimizer.py, dataset_training_service.py |
| 5 | Gradient clipping | derived+research-mode-only | MASS bounds updates in research modes; canonical AdamW path has no geometric clipper | mass_step_size.py |
| 6 | Warmup | derived+research-mode-only | canonical path uses calibrated cosine from step 0; research modes use MASS ceilings | dataset_training_service.py, mass_step_size.py |
| 7 | LR schedule | derived+research-mode-only | default: cosine over 6 data-epochs; no-schedule MASS only in research modes | _mlx_training_adapter_train_mixin.py, mass_step_size.py |
| 8 | Batch size | derived+shipped-default | derived from gradient-noise scale, then reduced only for memory-safe micro-batching | DatasetTrainingService.train_from_dataset |
| 9 | Early stopping | derived+shipped-default | geometric certificate and measured validation-loss windows are wired into training | geometric_early_stopping.py, _mlx_training_adapter_train_mixin.py |
| 10 | LoRA scale | derived+shipped-default | adapter scale budget and saturation telemetry are enforced during training | geometric_lora.py, _mlx_training_adapter_train_mixin.py |
| 11 | LoRA rank | derived+shipped-default | per-module ranks derive from tail dimensions and rank-capacity samples | geometric_lora.py, DatasetTrainingService.build_training_plan |
| 12 | Target modules | derived+shipped-default | target surface derives from layer spectral geometry | select_target_modules, DatasetTrainingService.build_training_plan |
| 13 | Dropout | removed | derived dropout formula was deleted because no shipped training adapter consumed it | compute_geometric_dropout runtime references=0 |
| 14 | Weight init | removed | default init is PiSSA; the unshipped spectral-normalized helper was deleted | spectral_normalized_lora_init runtime references=0 |
| 15 | Residual scaling | removed | standalone residual-scaling formula was deleted because no shipped path consumed it | residual_scaling.py; runtime references=0 |
Research program and formulas: 15-Hyperparameter Research Program and Geometric Hyperparameter Rosetta Stone.
git clone https://github.com/Ethyros-AI/ModelCypher.git
cd ModelCypher
poetry install # Primary macOS/MLX path; Python 3.11+
poetry run mc --help # Verify CLI installBackend protocol and CPU-fallback development on Linux use optional extras:
poetry install --extras cuda # NVIDIA/CUDA backend
poetry install --extras jax # JAX backend; CPU is supported for CI/testsThe MLX path is the only end-to-end product surface today. CUDA and JAX share the geometry protocol, but model-loading and expert-CLI parity remain partial; see BACKEND-COMPARISON.md.
# Inspect a model's per-layer geometry
poetry run mc model info /path/to/model
# Build an observation bundle from one prompt
poetry run mc analyze capture --model /path/to/model --prompt "Explain geodesics."
# Build an observation bundle from a prompt family
poetry run mc analyze family --model /path/to/model --manifest data/probes/prompt_family_minimal_pairs.json
# Re-read an existing bundle and print the shared report view
poetry run mc analyze report --bundle /path/to/bundle
# Re-read the retained measurement-atlas family root or one atlas run
poetry run mc analyze report --bundle results/measurement_atlas
poetry run mc analyze report --bundle results/measurement_atlas/<run_id>
# Re-read the retained pipeline-validation family
poetry run mc analyze report --bundle results/pipeline_validation
# Layer-wise intrinsic dimension profile
poetry run mc analyze dimension-profile --model /path/to/model --samples 50
# LoRA adapter spectral analysis
poetry run mc analyze lora-svd /path/to/adapter --base /path/to/model
# Train a LoRA adapter after inspecting the model
poetry run mc train run --model /path/to/model --data /path/to/data.jsonl --output /path/to/adapter| Question | What retained artifacts show | Tag |
|---|---|---|
| Does the measurement layer exist as a real workflow? | Yes. mc analyze capture, mc analyze family, and mc analyze compare now emit observation bundles with machine-readable artifacts plus a short report |
[EMPIRICAL] |
| Canonical training path exists | mc train run is the shipped geometry-derived runtime path guarded by pipeline_gate_v1 |
[EMPIRICAL] |
| Does the retained 350M validation bundle close preservation? | No. The tracked pipeline validation report records structural pass 5/5, inference pass 3/5, all_pass = false |
[EMPIRICAL] |
| Does the retained evidence close "better than standard practice"? | No. results/nblora_vs_standard/ is retained as summary_only, and the retained single-seed LFM2-350M summary does not support superiority of nb_lora over the kept baselines |
[EMPIRICAL] |
| Does the 8B bundle close efficacy? | No. Despite the historical family name, the tracked G5 report records n_seeds=1 and failed cka_ok and degenerate_ok gates |
[EMPIRICAL: 1 seed] |
| Is quantization promising? | Yes as a measurement surface: the tracked quantization frontier report records PPL and degeneration improvement on all 3 retained models, but the frontier law is still open | [EMPIRICAL] |
| Hypothesis | Result | Tag |
|---|---|---|
| REINFORCE on 350M | Gradient orthogonal to CE; degradation monotonic with steps | [DISPROVEN] |
| SFT on reasoning traces | Format memorization: PPL drops, inference degrades | [DISPROVEN] |
| Pullback metric P = MM^T | P ≈ I throughout training (median deviation 0.001) | [DISPROVEN] |
| Stable rank predicts adapter rank | Pearson r = -0.51 vs tail_dims; measures different property | [DISPROVEN] |
| Constrained training (paired) | Constraints monotonically hurt | [DISPROVEN] |
We publish failures because intellectual honesty is not optional. Current training blockers and exit criteria live in MISSION.md and RESEARCH-ROADMAP.md.
mc analyze is organized around five canonical workflows:
capture: measure one prompt or prompt filefamily: run explicit minimal-pair or perturbation studiescompare: compare two targets on the same prompt familyreport: read an existing bundle and render the shared high-signal viewprobe: targeted probe and red-team workflows
Expert instruments remain directly callable when you want the underlying measurements without the bundle wrapper:
- geometry and trajectory:
reasoning-flow,geodesic-profile,entropy-trajectory,chain-profile,dimension-profile,jacobian-trace - probe and monitoring:
calibrate-safety,jailbreak-test,probe-redteam,probe-behavioral,circuit-breaker,entropy-pattern - diagnostics and adapter analysis:
lora-svd,benchmark,knowledge-type,curriculum-profile,crm-build,crm-compare
Full reference: CLI-REFERENCE.md
Hexagonal (ports-and-adapters) with strict domain boundaries:
- Core domain (
core/domain/) — pure geometry and math, zero framework imports - Use cases (
core/use_cases/) — orchestration, cannot import from adapters - Adapters (
adapters/) — HuggingFace Hub, filesystem, model loading - Backends — MLX (primary, Apple Silicon), CUDA, JAX behind a protocol interface
Core geometric computations use the framework-agnostic Backend protocol, and backend selection is automatic when an installed backend is available. End-to-end loaders and expert CLI workflows remain MLX-first.
| Document | What It Covers |
|---|---|
| Start Here | Installation, first observation bundle, and downstream training |
| Observation Bundles | PromptFamilyManifest, ObservationBundle, and ready-to-run perturbation manifests |
| Geometry Guide | Interpreting CKA, intrinsic dimension, curvature, and entropy measurements |
| Training Guide | Downstream adapter workflows and dataset preparation |
| CLI Reference | Workflow-first mc analyze plus expert command examples |
| Mission | Measurement-first mission and derived training standards |
| 15-Hyperparameter Research Program | Per-control runtime and evidence status for the derivation program |
| Glossary | 60+ term definitions |
| Architecture | Hexagonal architecture and domain boundaries |
| Bibliography | All cited papers with local reference PDFs |
| Paper | Status | Thesis |
|---|---|---|
| The Shape of Knowledge | [EMPIRICAL] | Knowledge has measurable geometric structure; inference is trajectory |
| Invariant Semantic Structure | [PROVEN: fitted probes only]; [CONJECTURAL] cross-model | A closed-form fit gives aligned CKA = 1 on its training probes; held-out and cross-model invariance remain open |
| Entropy Safety Signal | [CONJECTURAL] | Behavioral drift detection via entropy differentials |
| Cross-Architecture Transfer | [CONJECTURAL] | Knowledge transfer between model families via Procrustes alignment |
| ModelCypher Toolkit | [EMPIRICAL] | Implementation methodology and CLI design |
| The Semantic Highway | [EXPLORATORY] | Historical TwoNN profile from an unretained run; published-profile replication is pending under WS4.2 |
7,758 collected tests. This count is generated from pytest --collect-only; refresh it with poetry run python scripts/update_test_count.py --write.
Includes Hypothesis property-based tests for numerical invariants (CKA symmetry, spectral bounds, null-space orthogonality).
poetry install --extras jax # One-time local CI dependencies
./scripts/run_local_ci.sh # Full local CPU verification
poetry run pytest # Standard development run
HYPOTHESIS_PROFILE=full poetry run pytest # Full property-based testing| Platform | Backend | Status |
|---|---|---|
| macOS Apple Silicon | MLX | Primary end-to-end product path |
| Linux + NVIDIA GPU | CUDA (PyTorch) | Backend protocol available; loader/CLI parity partial |
| Linux + TPU/GPU | JAX | Backend protocol available; loader/CLI parity partial |
| Linux CPU (local CI/tests) | JAX | Engineering fallback, not an accelerated product path |
@software{kempf2026modelcypher,
author = {Kempf, Jason},
title = {ModelCypher: MLX-Native Measurement Workbench for LLMs},
year = {2026},
url = {https://github.com/Ethyros-AI/ModelCypher},
license = {AGPL-3.0}
}AGPL-3.0. See LICENSE.