Imported from arjver/ucsb-academic-planner (
AGENTS.md). Install upstream withnpx skills add arjver/ucsb-academic-planner. Copyright stays with the author.
AGENTS.md
This file is a lightweight handoff guide for future work in this repository. It should point to the source of truth rather than duplicate large amounts of project documentation.
Project Purpose
This repository is building an AI-powered UCSB academic planner.
The high-level product goal and current roadmap live in:
Current Status
The project is in the early pipeline-build stage.
What is already done:
- Python project scaffold created with
uv - Core dependencies installed
- Initial Pydantic domain schemas written
- Minimal package layout created for future LangGraph pipelines
- Deterministic UCSB API ingestion layer implemented
- Live UCSB API verification completed against the ingestion pipelines
- Prerequisite scraper/parser pipeline implemented with Playwright + LangGraph + Gemini
- Scraper deduplicates course rows by
course_idto avoid overwriting parses with later empty sections - Pipeline supports
target_course_idsfor focused live runs and configurablemodel_name/requests_per_minute - Validate -> repair LangGraph loop fires on unresolved known-course mentions (ERROR) and unknown-course references
_convert_logicflattens empty groups to null and single-child AND/OR groups to their child
- Scraper deduplicates course rows by
- Degree requirements PDF extraction pipeline implemented with
pdftotext+ LangGraph + Gemini- Deterministic PDF extraction segments the major page into requirement blocks and strips most right-column GE/sidebar bleed
- Per-block LangGraph parsing writes structured
ProgramRequirementsartifacts and raw block artifacts to disk - Live extraction has been run successfully against
major-reqs/engineering.pdfforComputer Science 2025-26 - Current known quality gaps are truncated sidebar text on some left-column lines and footnote normalization (
CMPSC 196B2,CMPSC 1921,CMPSC 1961, etc.)
- Student profile normalization pipeline implemented
- Loose user-provided planning input can be normalized into canonical
StudentProfilerecords program_idcan be inferred fromprogram_name+catalog_year- Course references are normalized into stable
SUBJECT-NUMBERIDs for completed, in-progress, and constraint lists
- Loose user-provided planning input can be normalized into canonical
- Optimizer input contract implemented
OptimizerInputmerges student profile, requirements, courses, offerings, quarters, prerequisites, and constraints- Builder computes lightweight requirement progress for explicit course and choice rules
- Unit/GPA/residency/free-text rules are preserved but currently marked
unknownfor downstream/manual evaluation
- First scheduler-demo assembly path implemented
FIRST_SCHEDULER_DEMO_SCOPElocks the first engine target toComputer Science 2025-26build_demo_optimizer_input(...)builds a demo-ready merged input from the real parsed CS requirements artifact plus normalized student/catalog data- A realistic end-to-end demo student fixture exists in
tests/fixtures/optimizer/demo_student_profile.json
- Raw and parsed prerequisite artifact persistence implemented
- Live prerequisite verification completed on a mixed CMPSC batch:
CMPSC-32,CMPSC-40,CMPSC-130A/B,CMPSC-162,CMPSC-178,CMPSC-190G(covers AND/OR, min-grade, concurrent enrollment, consent-of-instructor, and major restrictions) - LangSmith tracing wired up end-to-end; graph/node calls show up under the
UCSB Academic Plannerproject whenLANGSMITH_TRACING=true - README updated with a concise tech stack and roadmap
What is not done yet:
- No optimizer or scheduler exists yet
- No UI or conversational interface exists yet
- No persistence layer exists yet
- Candidate course map used by the parser is scoped to the currently scraped quarter+subject, so cross-subject prereqs (MATH, PSTAT, ECE, ...) and courses not offered that quarter currently land in
ambiguitiesrather than the logic tree - Requirement extraction still needs stronger line normalization, footnote cleanup, and validation/evaluation coverage before it is reliable enough to feed an optimizer
- Optimizer input progress for unit/GPA/residency/free-text rules is intentionally shallow for now; those rule types still need richer evaluation logic before scheduling decisions can fully rely on them
Where Things Live
Core package:
Domain schemas:
- src/academic_planner/schemas/common.py
- src/academic_planner/schemas/catalog.py
- src/academic_planner/schemas/prerequisites.py
- src/academic_planner/schemas/requirements.py
- src/academic_planner/schemas/student.py
- src/academic_planner/schemas/plan.py
- src/academic_planner/schemas/extraction.py
- src/academic_planner/schemas/init.py
Pipeline scaffolding:
- src/academic_planner/graphs
- src/academic_planner/graphs/state.py
- src/academic_planner/graphs/prerequisite_state.py
- src/academic_planner/graphs/prerequisite_graph.py
- src/academic_planner/graphs/requirements_state.py
- src/academic_planner/graphs/requirements_graph.py
- src/academic_planner/nodes
- src/academic_planner/nodes/prerequisites.py
- src/academic_planner/nodes/requirements.py
- src/academic_planner/pipelines
- src/academic_planner/pipelines/prerequisite_pipeline.py
- src/academic_planner/pipelines/requirements_pipeline.py
- src/academic_planner/pipelines/student_profile_pipeline.py
- src/academic_planner/prompts
Prerequisite pipeline implementation:
- src/academic_planner/prerequisites
- src/academic_planner/prerequisites/scraper.py
- src/academic_planner/prerequisites/pipeline.py
- src/academic_planner/prerequisites/validation.py
- src/academic_planner/prerequisites/course_matching.py
- src/academic_planner/prerequisites/normalize.py
- src/academic_planner/prerequisites/artifacts.py
Requirements pipeline implementation:
- src/academic_planner/requirements
- src/academic_planner/requirements/pdf.py
- src/academic_planner/requirements/artifacts.py
- src/academic_planner/requirements/models.py
- src/academic_planner/schemas/requirements_artifacts.py
Student profile + optimizer input:
- src/academic_planner/students
- src/academic_planner/students/normalize.py
- src/academic_planner/optimizer
- src/academic_planner/optimizer/input.py
- src/academic_planner/optimizer/demo.py
- src/academic_planner/schemas/optimizer.py
Deterministic API ingestion:
- src/academic_planner/ingestion
- src/academic_planner/ingestion/raw_models.py
- src/academic_planner/ingestion/mappers.py
- src/academic_planner/ingestion/pipelines.py
- src/academic_planner/ingestion/ucsb_api.py
Project config:
Input reference material already checked into the repo:
Tests folder:
- tests
- tests/test_api_ingestion.py
- tests/test_prerequisite_pipeline.py
- tests/test_requirements_pipeline.py
- tests/test_optimizer_input.py
- tests/fixtures/optimizer/demo_student_profile.json
Source Of Truth
Use these as the canonical references:
- Product scope, tech stack, roadmap: README.md
- Python dependencies and tooling: pyproject.toml
- Current domain model contracts: src/academic_planner/schemas
- Current deterministic ingestion implementation: src/academic_planner/ingestion
- Current prerequisite scraper/parser implementation: src/academic_planner/prerequisites
- Current degree requirements PDF extraction implementation: src/academic_planner/requirements
- Current student normalization and optimizer-input implementation: src/academic_planner/students and src/academic_planner/optimizer
If there is ever a mismatch between this file and the code, trust the code and README first.
Working Conventions
- Keep the domain schemas as the stable contract between AI extraction and later optimization.
- Keep structured API ingestion deterministic; do not route UCSB API payloads through an LLM.
- Prefer adding or refining models in
src/academic_planner/schemas/before building pipeline logic around them. - Keep workflow state separate from business/domain models.
- Keep raw prerequisite scrape artifacts and parsed artifacts on disk for inspection while the pipeline evolves.
- Keep raw requirement block artifacts and parsed requirement artifacts on disk for inspection while the pipeline evolves.
- Avoid duplicating project planning notes across files; update the README roadmap instead.
- When adding a new major subsystem, create the code first and then add a short pointer here if needed.
- Store local secrets in
.envand never commit real keys.
Recommended Next Step
Start with the next unchecked roadmap item in the README:
- Build the scheduling engine that computes the most optimized academic plan given constraints
The prerequisite pipeline is already implemented and validated against a mixed live CMPSC batch.
The requirements pipeline now exists and has been live-run for Computer Science 2025-26.
It follows the same general reference pattern:
- deterministic scrape first
- structured LangGraph parse second
- validation and repair loop driven by
IssueSeverity.ERROR - post-parse normalization (flattening degenerate groups)
- artifact persistence for debugging
- LangSmith tracing via the
run_name/tags/metadataconfig on each LLM call
The most immediate known requirement-extraction issues after the first live run are:
- partially clipped right-column text can still land in course notes on the left column
- wrapped list lines sometimes lose a few trailing characters
- footnote-bearing course tokens like
CMPSC 196B2,CMPSC 1921, andCMPSC 1961need deterministic normalization before the model sees them
For a quick planner demo, the newly added optimizer input layer gives the scheduler a concrete target shape:
- normalized
StudentProfile ProgramRequirements- prerequisite/enrollment constraint data
- course catalog, offerings, and quarter timeline
- lightweight requirement progress summaries for explicit course and choice rules
The immediate handoff into the scheduler is now:
- use
FIRST_SCHEDULER_DEMO_SCOPE - use
build_demo_optimizer_input(...) - target the realistic CS scenario in
tests/fixtures/optimizer/demo_student_profile.json
When expanding the prerequisite parser further, the main known quality gap is still the candidate-course-map scope noted in Current Status; widening it beyond the currently-scraped quarter+subject would unlock cross-subject and not-offered-this-quarter prereq edges.
Quick Start
Set up and run Python commands from the repo root:
uv sync --extra dev
uv run pytest
uv run python -c "from academic_planner.ingestion import CurriculumIngestionPipeline; print('ok')"
uv run python -c "from academic_planner.pipelines import run_prerequisite_pipeline; print('ok')"
uv run python -c "from academic_planner.pipelines import run_requirements_pipeline; print('ok')"
uv run python -c "from academic_planner.pipelines import run_student_profile_pipeline; print('ok')"
Conversation Handoff
When starting a new conversation, read these first:
- README.md
- AGENTS.md
- src/academic_planner/schemas
- src/academic_planner/ingestion
- src/academic_planner/prerequisites
- src/academic_planner/requirements
That should be enough context to continue work without re-planning the project from scratch.