Imported from hittanshubhanderi20/stitchfam (
AGENTS.md). Install upstream withnpx skills add hittanshubhanderi20/stitchfam. Copyright stays with the author.
AGENTS.md — Cross-Family Stitching Project
What this project is
Research project testing whether stitching decoder blocks from LLMs with different tokenizers (e.g. Qwen + Gemma) works, and how performance changes as the number of stitch points increases. Full step-by-step plan lives in the user request.
Before doing anything
- Read
PROGRESS.mdin full. It records exactly which numbered steps are done, blocked, or not started, and any decisions/deviations from earlier sessions. - Do not start a step whose prerequisite step is not marked "done" in
PROGRESS.md. - If
PROGRESS.mdshows a step as "blocked," resolve or explicitly re-scope the blocker before proceeding — do not silently skip it.
While working
- Work strictly one numbered step at a time, and never proceed to or execute more than one step at a time under any circumstances.
- A step is only "done" when its concrete output exists on disk (a file, a result, a checkpoint) — not merely attempted.
- Keep the adapter architecture and seam-placement rules fixed across the Phase 9 sweep — do not vary more than one thing per experiment.
- If there is any step or action that cannot be done directly by you, explicitly tell the user and they will do it for you.
- At the end of every step, verify that everything is working as expected and has completed successfully before updating
PROGRESS.md.
After completing each step
Immediately append one entry to PROGRESS.md (do not batch multiple steps into one entry, and do not wait until the end of the session). Use the exact format defined in PROGRESS.md's header.
Compute rules
- Kaggle is the default compute engine for anything requiring backprop or GPU inference — never the local M2 Pro.
- Move a run to Colab Pro only if it would exceed Kaggle's session length, exceed the remaining weekly Kaggle quota, or need more VRAM than Kaggle provides. Log which platform each run used in PROGRESS.md.
- Data prep, orchestration, analysis, and plotting run locally.
Step-by-Step Plan
Phase 0 — Environment & Project Scaffolding
- Create the project's root folder locally, named e.g.
stitchfam/. (Done) - Inside it, create the subfolder structure:
paper/,notes/,src/,data/,results/,checkpoints/,README.md. (Done) - Initialize a git repository inside
stitchfam/(git init). - Create a
.gitignorethat excludescheckpoints/,data/, and any.envfiles. - Create a private GitHub repo and push the empty scaffold, so history is tracked from day one.
- Create a Python virtual environment on the M2 Pro (
python -m venv .venv) for local work (data prep, zero-shot experiments, analysis). - Install local-only dependencies:
transformers,torch(CPU/MPS build),numpy,pandas,matplotlib. - Set up a Kaggle account and confirm GPU access (P100 or 2×T4) — this is the default compute engine for the project. Note the weekly 30-hour quota and ~9-12 hour session limit in
notes/compute.md. - Set up Google Colab Pro as the overflow compute environment — confirm access to A100 (default GPU choice for this project) and note in
notes/compute.mdthat H100 is used only as a fallback if A100 availability becomes a bottleneck, since this project's compute needs don't require H100. Note the available compute unit balance and treat it as fully available to spend on this project rather than something to conserve. - Write the allocation rule into
notes/compute.md: Kaggle is used by default for every GPU step in this plan; a step moves to Colab Pro only when it would exceed Kaggle's session length, exceed the remaining weekly Kaggle quota, or need more VRAM than Kaggle's P100/2×T4 provides. - On both Kaggle and Colab notebooks, install the training-side dependencies at the top of every session:
transformers,accelerate,bitsandbytes,datasets,wandb,huggingface_hub(each platform's owntorch/CUDA build is already present — don't reinstall it). - Create a free Weights & Biases account and connect it (
wandb login) for experiment tracking — this matters more with two compute platforms in play, since it gives you one unified run history regardless of whether a given run happened on Kaggle or Colab. - Create a private Hugging Face Hub repo (or repos) to serve as your persistent storage for anything you want to keep across sessions and across platforms — trained adapter checkpoints, primarily. This is what makes switching between Kaggle and Colab mid-project seamless, since neither platform's own storage is the source of truth.
- Mount Google Drive in Colab notebooks (and use Kaggle's own persistent notebook storage on Kaggle) only for lightweight persistence:
results/sweep.csv, plots, andnotes//PROGRESS.md— never for full base model weights. - Note in
notes/compute.md: base model weights (Qwen, Gemma, Phi, etc.) are downloaded fresh from Hugging Face Hub to whichever platform's local disk is in use at the start of every session, since both Kaggle's and Colab's local disks are ephemeral — never assume a previously downloaded model is still there when a new session starts. - Write a one-page
notes/research_question.mdstating the hypothesis precisely (as written at the top of this document) so it's pinned down before any code is written. - Set up the Antigravity IDE project files described in Part 2 (
AGENTS.md,PROGRESS.md, and the/complete-stepworkflow) before starting Phase 1.
Phase 1 — Grounding: Read Before You Build
- Download the StitchLLM paper PDF into
paper/stitchllm.pdf. - Download the Relative Representations paper (Moschella et al., 2023) into
paper/relative_representations.pdf. - Download the "Cross-Tokenizer LLM Distillation through a Byte-Level Interface" paper into
paper/cross_tokenizer_distillation.pdf. - Write a one-paragraph summary of each paper's method and result in
notes/summary.md, in your own words, before using AI to assist further. - Clone
lucmos/relrepsinto athird_party/folder — this is your zero-shot stitching baseline implementation. - Clone
arcee-ai/mergekitintothird_party/— this is your same-family control baseline tool. - Clone
EleutherAI/lm-evaluation-harnessintothird_party/— this will run your benchmarks (MMLU, etc.) so you don't hand-roll evaluation code.
Phase 2 — Model Selection
- Choose your primary cross-family pair: a larger "donor" model and a smaller "receiver" model, from different tokenizer families. Recommended starting pair: Qwen2.5-3B (donor, bottom layers) and Gemma-2-2B or Gemma-1B (receiver, top layers).
- Choose a second, held-out pair for cross-validation of your findings — e.g. Phi-3.5-mini (donor) and a small Qwen model (receiver).
- Confirm both models' tokenizers are genuinely different (different vocab size, different merge rules) — write this confirmation into
notes/model_pairs.mdwith each model's exact HuggingFace ID and tokenizer vocab size. - Download both models in your primary pair from HuggingFace onto the cloud GPU instance (not the laptop).
- Run a basic forward pass of each model individually on a test sentence, on the cloud GPU, to confirm both load and generate correctly before any stitching work begins.
- Record each standalone model's baseline benchmark score — this is the reference number every stitched result gets compared against.
Phase 3 — Calibration Data
- Choose a calibration text corpus for collecting paired activations — a few thousand general-domain sentences (e.g. a WikiText-103 subset) is sufficient at this scale.
- Download and save the calibration corpus into
data/calibration/. - Write a script
src/extract_activations.pythat runs a batch of calibration text through a given model and saves the hidden state at every decoder layer to disk. - Run
extract_activations.pyon the donor model (Qwen2.5-3B), saving all layer activations for the calibration set. - Run
extract_activations.pyon the receiver model (Gemma), saving all layer activations for the same calibration set. - Verify both activation dumps have the same number of samples in the same order — write a small assertion script for this.
Phase 4 — Baseline 1: Zero-Shot Stitching (No Training)
- Set up the
relrepscodebase to compute relative representations (anchor-based cosine similarity encodings) for the donor model's activations. - Do the same for the receiver model's activations, using the same anchor set drawn from the calibration data.
- Implement the zero-shot stitch: replace the receiver model's early-layer activations with the donor's relative-representation-transformed activations at your chosen single splice point.
- Run this zero-shot stitched pipeline on a small held-out text sample and manually inspect the generated output for basic coherence.
- Record the zero-shot stitching result as Baseline 1 in
results/baseline1_zeroshot.md— this establishes your performance floor before any adapter training.
Phase 5 — Baseline 2: Same-Family Control
- Pick two same-family models of different sizes (e.g. two Qwen2.5 sizes) as your "easy mode" control group.
- Use
mergekit-moeto upcycle these two same-family models into a merged model, following the tool's standard workflow. - Run this same-family merged model through the same benchmark you'll use for the cross-family experiments (defined in Phase 6).
- Record this as Baseline 2 in
results/baseline2_samefamily.md— this tells you the "ceiling" of what stitching/merging can achieve when tokenizer compatibility isn't a problem.
Phase 6 — Benchmark Setup
- Choose one primary benchmark for all experiments — start with a manageable MMLU subset (a handful of categories, not the full suite, to keep GPU time reasonable).
- Configure
lm-evaluation-harnessto run that MMLU subset against a HuggingFace-format model. - Do a dry run of the harness against one of your unmodified standalone models to confirm the evaluation pipeline works end-to-end before touching stitched models.
- Decide and document your single "result" metric precisely — this is the number every configuration in your sweep will report.
Phase 7 — Trained Single-Point Stitching Adapter
- Implement a stitching adapter class in
src/stitch_adapter.py— start with the simplest version: a single linear (affine) layer mapping donor hidden-dim → receiver hidden-dim. - Write the training loop: freeze both donor and receiver model weights entirely, and only update the adapter's parameters.
- Train this single-point adapter using the paired activations from Phase 3, using StitchLLM's reported scale as your budget target (under ~2,000 gradient steps).
- Save the trained adapter checkpoint to
checkpoints/single_point_v1.pt. - Run the full stitched pipeline (donor bottom → adapter → receiver top) through the Phase 6 benchmark.
- Record this as Result: 1 stitch point in
results/sweep.csv, alongside the two baselines. - Compare this trained result against Baseline 1 (zero-shot) — confirm training actually helped before proceeding.
Phase 8 — Multi-Point Stitching Infrastructure
- Extend
stitch_adapter.pyto support inserting more than one adapter at more than one layer boundary in the same forward pass (not just one bottom+top split). - Decide your seam placement strategy for multi-point runs — e.g. evenly spaced through depth — and keep this rule fixed across the whole sweep so seam count is the only thing changing.
- Keep the adapter architecture (a single linear layer) identical at every seam for every configuration in the sweep — do not let adapter capacity change alongside seam count.
- Implement a config file (
configs/sweep.yaml) listing each stitch-count configuration to run (e.g. 1, 2, 4, 8 seams) and their fixed seam positions.
Phase 9 — Run the Stitch-Count Sweep
- Train the 2-seam configuration adapter set, following the same training budget as Phase 7.
- Benchmark the 2-seam model and append the result to
results/sweep.csv. - Train the 4-seam configuration adapter set.
- Benchmark the 4-seam model and append the result to
results/sweep.csv. - Train the 8-seam configuration adapter set.
- Benchmark the 8-seam model and append the result to
results/sweep.csv. - (Stretch) Train and benchmark a 16-seam configuration if the 8-seam result is still trending upward and GPU budget allows.
- Repeat steps 61–67 for your second, held-out model pair to check whether the curve's shape generalizes across pairs.
Phase 10 — Sanity Checks Against False Positives
- Run the "stitch hacking" check: feed random noise through the donor model's portion of the pipeline instead of real activations, and re-run the benchmark. If performance doesn't collapse, the receiver is likely ignoring the donor and coasting on its own capability — flag any configuration where this happens.
- For any configuration flagged in step 69, note it explicitly in your results as unreliable rather than silently including it in your headline numbers.
- Spot-check generated text quality manually for the best and worst configurations in the sweep, to catch cases where a high score doesn't correspond to sensible output.
Phase 11 — Analysis
- Plot the full sweep as a single chart: x-axis = number of stitch points, y-axis = benchmark score, one line per model pair, both baselines shown as horizontal reference lines.
- Write up, in your own words, whether the curve is monotonic, plateaus, or reverses past some seam count — and propose a reason why, grounded in what you know about error compounding across seams.
- Compare your cross-family sweep curve directly against Baseline 2 (same-family) to quantify the concrete cost of crossing tokenizer families.
- Write a short "Limitations" section honestly noting anything from step 69/70 that affects how much to trust the headline results.
Phase 12 — Publish
- Write the final
README.mdcovering: project overview, links to the three grounding papers, your exact model pairs, your sweep results table and chart, comparison against both baselines, what you learned, and what you'd try next. - Push the final repo, including
results/sweep.csvand the chart image, to GitHub. - Write a short (one paragraph) plain-language summary of the finding for your resume/portfolio summary.
- Once the above is done and the project's findings are locked in, upgrade to Hugging Face PRO ($9/month) — this step is deliberately last.
- Build a small Gradio demo app that loads your best-performing stitched configuration and lets a visitor enter a prompt and see the output.
- Deploy the demo as a Hugging Face Space with ZeroGPU enabled, using the PRO tier's higher daily quota and H200 access.
- Link the live Space demo from the GitHub
README.mdand from your resume/portfolio summary.
