Imported from Younesb24/pfa-decision-ai (
docs/agent_docs/AGENTS.md). Install upstream withnpx skills add Younesb24/pfa-decision-ai --skill agent_docs. Copyright stays with the author.
PFA — Outil IA pour l'Aide à la Décision en Entreprise
Context
- PFA MGSI ENSAO 2026 — AI-powered decision support for e-commerce marketplace operations
- Dataset: Olist Brazilian E-Commerce (8 relational CSV tables, ~100K orders)
- Solo student project, 4-month timeline, must ship a working MVP
Persona
- Responsable Opérations E-Commerce (Head of E-commerce Ops)
- Needs: weekly operational briefing, KPI monitoring, anomaly alerts, delivery risk management
Stack (LOCKED — do not suggest alternatives)
- Storage: PostgreSQL 15-alpine (Docker)
- Analytics: DuckDB only for ad-hoc notebooks (NOT a runtime layer; the API reads Postgres exclusively — see ADR-005)
- Transform: dbt Core 1.11 (staging → marts)
- Backend: FastAPI + Pydantic v2 + raw psycopg2 (no ORM — we own our SQL)
- Frontend: Next.js 16.2.6 App Router + Shadcn/UI + Recharts + TypeScript strict
- LLM: OpenAI GPT-4o (the configured provider for this deployment) with Anthropic Claude Sonnet supported as a drop-in alternative if ANTHROPIC_API_KEY is set. Used for narrative generation + the tool-using agent. NEVER raw text-to-SQL on bronze.
- ML: scikit-learn + XGBoost + isotonic calibration
- Auth: JWT + bcrypt; RBAC enforced on write endpoints; read endpoints are MVP-grade (not gated)
- Orchestration: Dagster (replay clock + nightly sensors + dbt scheduling)
- Container: Docker Compose
- CI/CD: GitHub Actions
Non-Negotiable Rules
- LLM = narrator on pre-computed KPIs, NEVER calculator on raw data
- Bronze raw data NEVER exposed to frontend or LLM — only validated Gold layer
- Every LLM recommendation must display its source data
- Deterministic computations separated from narrative layer
- All KPIs defined in
docs/agent_docs/kpi_catalog.md— check before creating new ones
Code Standards
- Python 3.11+, type hints mandatory, ruff + mypy
- dbt models: snake_case, layer prefixes (stg_/int_/fct_/dim_/agg_)
- SQL: PostgreSQL dialect. Application SQL (api/routers/*.py) never uses
SELECT *— every column explicit. dbt marts are allowed to useSELECT *in CTEs (idiomatic dbt pattern) but the final model SELECT enumerates columns. - API: FastAPI + Pydantic v2 DTOs, dependency injection for DB, OpenAPI auto-doc
- Frontend: Next.js 16.2.6 App Router, Shadcn/UI, TypeScript strict, no
any - Tests: pytest for Python, dbt tests for models, Vitest for Next.js
- Notebooks: must be valid Jupyter JSON. CI lints
api ml scriptsonly; notebooks are not linted but MUST parse withjson.load.
NEVER Suggest (during MVP)
- Kafka, LangChain, Airflow, Spark (distributed)
- Multiple datasets (DataCo, Budget vs Actual) — Olist only
Future Add-ons (post-MVP, after ship)
- ChromaDB + RAG on Olist reviews (~99K PT comments)
- SpringBoot backend for ERP/CRM integration layer & Multi-tenant auth, enterprise SSO, mobile app
- Azure Data Factory + Databricks (cloud-native orchestration)
- Celery + Redis (async job queue)
- Snowflake free trial (cloud mirror of Gold layer)
- Self-Critique LLM (second pass verification)
- Sentiment NLP on Portuguese reviews
Definition of Done
- Code on feature/ branch, tests pass
- dbt tests green for any SQL changes
- OpenAPI docs updated for API changes
- README updated if needed
- No secrets in code (use .env.example)
