Imported from Zhaozed/Loona-PMdevelopskill (
skills/loona-test-project/SKILL.md). Install upstream withnpx skills add Zhaozed/Loona-PMdevelopskill --skill loona-test-project. Copyright stays with the author.
Loona Intent Testing
Use the bundled Rust tester for repeatable regression tests after every model training.
API
Default endpoint is configured in loona-intent-tester/src/config.rs:
http://117.50.221.35:35000/classify
Request: { "text": "..." }.
Build
cd /Users/joey/.codex/skills/loona-test-project/loona-intent-tester
cargo build --release
Binary:
/Users/joey/.codex/skills/loona-test-project/loona-intent-tester/target/release/loona-intent-tester
Standard full-language test
Use the final multilingual workbook directory. Start with --dry-run, then run the API test:
/Users/joey/.codex/skills/loona-test-project/loona-intent-tester/target/release/loona-intent-tester \
--input-dir "/path/to/corpora" \
--all-languages \
--output-dir "/path/to/corpora/test_reports" \
--concurrency 20 \
--threshold 0.85 \
--max-retries 3
For only formal test sheets, add repeated --include-sheet options. For one language use --language thai (aliases such as th and 泰语 are supported).
Modes
--dry-run: detect language/intent columns and count extracted tasks without API calls.--include-file TEXT: test only filenames containing TEXT; repeatable.--include-sheet NAME: test exact sheet names; repeatable.--sheet-contains TEXT: test sheets whose names contain TEXT; repeatable.--retry-api-failures-from REPORT.xlsx: retry only API failure rows.--resume CHECKPOINT.json: resume an interrupted run.--retest-included-sheets: retest selected sheets and merge into the latest language reports.--use-existing-reports: reuse existing latest reports when possible.--rebuild-existing-reports: rebuild report layout without API calls.
Report interpretation
Each row is classified as:
完全正确: predicted label matches true label and score meets threshold.正确但置信度不足.错误但置信度高: highest-priority model/label-boundary issue.错误且置信度低.API请求失败.
Reports are written as intent_<language>_<timestamp>.xlsx and intent_overall_<timestamp>.xlsx, containing summary, detailed results, per-sheet summaries, and error-category sheets.
Regression-test policy
For every newly trained model, run all supported languages on the same frozen test set. Track overall accuracy, per-language accuracy, per-intent recall, high-confidence error count, API failure rate, and latency. Compare against the previous report before approving the model.
Do not use the legacy scripts/test_loona_tables.py unless its endpoint and payload are updated; it contains an older API configuration.
