Imported from Zoh007/llm-judge (
AGENTS.md). Install upstream withnpx skills add Zoh007/llm-judge. Copyright stays with the author.
AGENTS.md
Cursor Cloud specific instructions
Project overview
LLM Judge — score coding LLM / student answers without execution using Sentence Transformers + logistic regression. Streamlit UI + CLI.
Virtual environment
source /workspace/.venv/bin/activate
# or from repo root:
# python -m venv .venv && source .venv/bin/activate && pip install -r requirements.txt
Running the application
Streamlit (from coding-llm-eval/):
cd /workspace/coding-llm-eval && streamlit run app_streamlit.py --server.port 8501 --server.headless true
CLI scoring:
cd /workspace/coding-llm-eval && python experiments/evaluate_csv.py \
--input_path examples/sample_submissions.csv \
--output_path examples/sample_scored.csv
Key caveats
- Training data is not committed. See
coding-llm-eval/data/README.md. - Trained models (
*.pkl) are gitignored; runpython experiments/train_judge.pyafter preparing data. - If a pickle fails to load (e.g. saved on Apple MPS), retrain locally.
- Shared scoring helpers live in
coding-llm-eval/core/(hosted API coming soon). - No linter or test framework is configured yet.
