Imported from yuyi2439/TokenDye (
AGENTS.md). Install upstream withnpx skills add yuyi2439/TokenDye. Copyright stays with the author.
TokenDye - Agent Guide
Project standards
- Core library:
src/tokendye/- write code and comments in English. - Docs:
README.md- write in English. - Library scope: the library only defines dye modules (
DyeLayer,DyeStack,DyeConfig) and inference-time injection (TokenDye). Training is NOT the library's concern — no training loops, losses, model loading, or RL machinery insrc/tokendye/. Those live inexample/(shared helpers inexample/common.py). - DyeLayer is pure: it carries no runtime state; label masks are passed
through
forward(hidden, labels)or held byDyeStack. Never add batch state toDyeLayer. - Published artifact: a dye ships as a pair —
dye_final.pt(weights) plusdye_config.json(labels, shape, and placement). Inference loads viaTokenDye.from_files(...).apply(model); see README for the exact usage. - Tooling: manage dependencies and run scripts with
uv; prefer latest dependency versions. Dependency layout (hfextra, per-examplerequirements.txt) and example commands are documented in README — keep them in sync there, don't duplicate them here. - No machine-specific paths: never hardcode absolute paths such as
/home/...in docs, docstrings, or code (the repo is published). - Linting: use
ruff(dev dependency group); BLE001 (blindexcept) is intentionally ignored. - No logging in the library:
src/tokendye/must not contain any logging-related code. Dev scripts log directly withloguru(real-time stderr + file sinks). - No domain rewards in the library: reward/judge logic is task-specific;
keep it in
example/Qwen3.5/(e.g.judge.py), not insrc/tokendye/. - RL rollout memory: generate with
past_key_values(KV cache); a full-sequence recompute per token OOMs at long lengths (Qwen3.5 also materializes seq^2 attention in its linear-attention torch fallback). - Data validation: malformed training/attack data must raise an error before any training starts (never train silently on empty targets).
Design and usage
See README.md for everything else: dye-method designs (including
the 1 + delta parameterization and fp32 master weights), how to run the
Qwen3.5/RWKV7 examples (model paths, env vars, commands), the training
objective, data formats, and direct package usage.
