Imported from zkf516/THU-BDC2026 (
AGENTS.md). Install upstream withnpx skills add zkf516/THU-BDC2026. Copyright stays with the author.
Repository Guidelines
Project Structure & Module Organization
This is a Python baseline for the 2026 THU Big Data Competition stock-ranking task. Core code is in code/src/: config.py centralizes parameters, model.py defines StockTransformer, utils.py handles feature engineering, train.py trains, and predict.py writes predictions. Data is under data/, artifacts under model/, predictions under output/, temporary scoring files under temp/, docs images under asset/, and verification scripts under test/.
Build, Test, and Development Commands
uv sync: install Python dependencies frompyproject.tomlanduv.lock.source .venv/bin/activate: activate the environment on Linux/macOS.sh train.sh: train and save artifacts undermodel/.sh test.sh: runpython code/src/predict.pyand generateoutput/result.csv.python test/score_self.py: scoreoutput/result.csvagainstdata/test.csv.docker buildx build --platform linux/amd64 --build-arg IMAGE_NAME=nvidia/cuda -t bdc2026 .: build the submission image.docker compose up: validate the packaged runtime and expectedtest/output/result.csv.
Coding Style & Naming Conventions
Use Python 3.10-3.12, 4-space indentation, snake_case functions and variables, and PascalCase classes. Keep configuration in code/src/config.py and paths relative to the repository root. Preserve Chinese data columns such as 股票代码, 日期, and 开盘. No formatter or linter is configured; keep imports tidy.
Testing Guidelines
No formal unit-test framework is configured. Use python test/score_self.py for scoring logic, sh test.sh for prediction output, and sh train.sh when training changes. For Docker or submission changes, run docker compose up. Always inspect the generated CSV.
Competition Rules & Output Contract
The task is to predict a CSI 300 portfolio with the highest next-week return. result.csv must be UTF-8 text with columns stock_id,weight. Submit at most five distinct stock codes. Weights must be numeric and sum to no more than 1; unused weight is cash. Evaluation buys at T+1 open and sells at T+5 open. Avoid wrong column names, duplicate IDs, more than five rows, non-numeric weights, empty/binary files, Chinese comma separators, or sums above 1.
Results must come from real training and prediction. Do not hard-code final stock codes or weights. Reproduction must regenerate the same result.csv from submitted code and input data.
Final validation mounts data/, output/, and temp/, overwriting image contents. Keep custom data or durable artifacts under model/ or another non-mounted directory. External data must be freely public and documented. Pretrained models must use non-commercial public data or have been publicly released before 2026-04-01.
Commit & Pull Request Guidelines
Use Conventional Commits, for example feat(model): add hist gradient boosting classifier, test(model): record local score, or chore(result): update prediction output. Keep commits focused and split code, tests or scoring records, generated results, and model documentation into separate commits.
Prediction-layer changes belong on main. New model experiments must start from the current main branch, use a model-name branch such as model/hist-gradient-boosting, and include a Chinese Markdown record under the matching model/ subdirectory. Pull requests should include motivation, changed pipeline stages, commands run, score or output path, reproducibility notes, and data/model artifact assumptions.
Security & Configuration Tips
Large CSVs, trained weights, and TensorBoard logs can be expensive to regenerate; do not replace them casually. get_stock_data.py depends on Baostock availability and date settings, so document date-range changes.
