Imported from YapWH1208/llm-evaluation (
AGENTS.md). Install upstream withnpx skills add YapWH1208/llm-evaluation. Copyright stays with the author.
Project guidance
LLM/SLM evaluation platform: FastAPI backend (backend/, Python ≥3.12, SQLite-first with an optional MongoDB document store) + React 19 / Vite / TypeScript web app (frontend/). End-to-end flow: docs/evaluation-workflow.md. Ops: docs/deployment.md. Human-facing developer guide: DEVELOPER.md.
Skills — STOP. Load a skill before you act.
Skills (playbooks) are user-level, in ~/.agents/skills/, loaded via the skill tool — never run_skill, and the names have no superpowers- prefix. A repo-local skill (superdesign) may live in .agents/skills/ (gitignored, pinned via skills-lock.json).
The rule: before you do ANYTHING non-trivial — before you explore, run bash,
write code, or answer — STOP and check the skills index. If a skill might fit, load it.
Loading a skill is cheap. Skipping it is the #1 mistake. Process skill FIRST. Action SECOND.
| If… | Load this FIRST |
|---|---|
| starting a feature, or you have a rough idea | brainstorming |
| a bug, a failing or flaky test, or anything surprising | systematic-debugging |
| writing or fixing any code | test-driven-development |
| you have a spec for a multi-step task | writing-plans |
| executing a written plan in this session | executing-plans |
| about to say "done" / "fixed" / "passing" | verification-before-completion |
| work is done and tests pass | finishing-a-development-branch |
you got code-review feedback (from review or a human) |
receiving-code-review |
| need an isolated workspace | using-git-worktrees |
| making or editing a skill | writing-skills |
These skills supplement the native tools — they don't replace them. For
dispatching subagents, code review, parallel work, and codebase exploration, use the
native tools directly: task (run a subagent), review (code-review a diff),
wait (join parallel jobs), explore (investigate the codebase).
If you catch yourself about to explore, fix, or answer without loading a skill — STOP and load it. "Build X" → brainstorming first. "Fix this bug" → systematic-debugging first.
Commands (repo root)
- Install backend deps:
python -m pip install -e ".[dev]"(editable install from repo root;pythonpath=["backend"]comes from pyproject.toml). No requirements.txt. MongoDB needs the extra:"…[dev,mongodb]"(pymongo is not in the base dev install). - Backend tests:
python -m pytest -q(config setspythonpath=["backend"],testpaths=["tests"]). Single file:python -m pytest tests/test_x.py -q. - Run API:
uvicorn app.main:app --app-dir backend --reload(docs at/docson port 8000). - DB CLI:
python -m app.cli database preview|initialize|validate(migrations also run automatically at startup). - Docker deploy:
docker compose up --build(API with SQLite + web container) needs no environment variables — the container auto-provisions a Fernet key at./data/.lle-secret-keyon first start; see docs/deployment.md. The API has no built-in authentication — keep it on loopback or a trusted network. - Frontend deps/dev server:
npm ci/npm run devinsidefrontend/(vite, port 5173). Rootpackage-lock.jsonis a stub — nevernpm installat repo root. Node 22 is what CI uses. - Frontend tests:
npm test -- --runinfrontend/— barenpm testis vitest watch mode (config lives invite.config.ts: jsdom,src/test-setup.ts). - Frontend typecheck + build:
npm run build(runstsc -b && vite build). No lint script exists for either end. - One-command local launch:
quick-launch.sh/quick-launch.command/quick-launch.bat(installs deps, starts both services, createsdata/.lle-secret-keyon first run).
Environment (LLE_*)
LLE_SECRET_ENCRYPTION_KEY— Fernet key for encrypted endpoint credentials; if unset the launchers persist one atdata/.lle-secret-key(gitignored).LLE_DATABASE_URL— defaults tosqlite:///./data/llm_evaluation.db;mongodb://switches to the document store.LLE_DATABASE_INIT_MODE—auto_migrate(default) |preview|validate.- Full list (concurrency ceilings, dataset host allowlists, CORS, etc.):
backend/app/core/config.pyis the single source of truth.
Architecture
- App factory
create_app()inbackend/app/main.py; feature routers, services, and ports live inbackend/app/modules/*; SQLAlchemy models and forward-only migrations remain underbackend/app/db/. - Application behavior is storage-neutral. SQLite and MongoDB differences live in repository adapters selected once by
create_app(); feature APIs and services must not branch on the configured backend. - Provider protocols have one adapter each under
backend/app/infrastructure/providers/adapters/. Built-in benchmark plugins live inbackend/app/benchmarks/; their persisted catalog is owned bybackend/app/modules/benchmarks/. - Frontend transport is limited to
frontend/src/shared/api/client.tsanderrors.ts; feature API operations, DTOs, state, and effects live underfrontend/src/features/*.
Testing quirks
- MongoDB tests fake the client — never require a live MongoDB.
- Avoid wall-clock-dependent tests (history: rate-window tests were made deterministic for this reason).
- CI (
.github/workflows/ci.yml) runs Ruff lint/format andpython -m pytest -q, then infrontend/:npm ci, ESLint, tests, and the production build.
Frontend conventions
- All UI copy goes through the typed i18n catalog
frontend/src/i18n/catalog.ts(8 locales, strict key parity — a missing key failssrc/i18n/locales.test.ts). - Tests are colocated as
src/*.test.ts(x)(vitest + jsdom + Testing Library).
Repo gotchas
.code-review-graph/is a local-only knowledge graph DB (gitignored). A pre-commit hook runscode-review-graph updateanddetect-changes --brief. If the graph warns it was built on another branch, runcode-review-graph build./docsand/dataare gitignored (existing docs are force-tracked; new files underdocs/needgit add -f).reasonix.tomlis gitignored too.- Default branch is
master; feature work is done onagent/...branches merged via PRs. Keep commits conventional (feat:,fix:,docs:,test:,ci:). - Keep
CHANGELOG.mdupdated under## Unreleased(or a new version section) when user-facing behavior changes.
