Imported from el-noir/covenant_sentinel (
AGENTS.md). Install upstream withnpx skills add el-noir/covenant_sentinel. Copyright stays with the author.
Covenant Sentinel
A financial covenant-monitoring agent built for the Vultr Cloud x 24 Raise Summit (Statement Two track). It accepts real credit agreement PDFs and pre-loaded SEC documents, extracts covenant clauses via LLM, calculates financial ratios via genuine LLM tool-calling on Vultr kimi-k2.6, checks historical trends, makes a preliminary decision, then challenges and reconciles its own findings before producing a confidence-scored memo.
Pitch (One-Liner)
"Covenant Sentinel reads your actual credit agreement, extracts the covenants, and argues with itself before escalating. Finance teams stop drowning in false-positive alerts."
Hackathon Constraints (HARD RULES)
Do NOT build any of these — they result in immediate disqualification:
- Streamlit applications — If we ship Streamlit, we are disqualified.
- Basic RAG applications — A single retrieve-then-answer pipeline is banned.
- Dashboard as the main feature — Passive data viewers are banned.
- Presentations, not technical demos — Judges want screen recordings of actual software, not slides.
What This Means for Architecture
- Next.js is mandatory, not optional. We ship a real web app with a build step.
- The frontend must show the agent reasoning, not just the result. The
StepTracecomponent that visualizesplan → retrieve → calc → decide → challenge → reconcile → memois not a decoration — it is the primary feature. It proves this is an agent, not a dashboard. - The adversarial layer is non-negotiable. Without
challenge → reconcile, this is basic RAG + a flag. The self-checking loop is what transforms a retrieval pipeline into an agent. - Demo video = 50% of the score. Every architectural decision must optimize for what is visible in a 60-second screen recording.
Data Input Model — Dual Mode
The agent supports two ways to provide documents:
- Pre-loaded Collection (Default for demo): User selects a real company (MosaicCo, CovestroAG). Data comes from
data/covenants.json— real SEC credit agreement covenant sections with representative legal language. - PDF Upload (Feature): User uploads any credit agreement PDF. The backend parses text with
PyPDF2, uses an LLM to extract covenant clauses / thresholds / formulas, validates the extraction, stores it temporarily, then runs the full analysis.
Demo strategy: The 60-second video only uses pre-loaded mode to guarantee stability. The upload tab is visible in the UI as proof of capability but is not triggered live unless a judge asks during Q&A.
Architecture
State Schema (TypedDict)
class AgentState(TypedDict):
company_id: str
covenant_type: str # "fixed_charge_coverage" | "leverage_ratio"
plan: str # LLM reasoning for selected covenant
clause_text: str # Full clause text retrieved
clause_citation: str # e.g. "Section 6.1, Credit Agreement dated March 14, 2024"
financial_inputs: dict # {"ebitda": 1200000, "fixed_charges": 810000}
trend_values: list[float] # Last 4 quarters
ratio_value: float
threshold: float
ratio_formula: str
preliminary_decision: str # "BREACH" | "NO_BREACH" | "WATCH"
preliminary_confidence: float
challenge_arguments: str # Devil's advocate output
final_decision: str
final_confidence: float
memo: str
_loop_count: int # Guards adversarial cycles (max 1)
Graph Flow
The agent is a ReAct-style multi-step loop. At key decision points (calc_agent and memo_agent), the LLM is bound to a shared tool pool. If it decides a tool is needed, execution routes to a single ToolNode, which executes the chosen tool and returns the result.
plan -> retrieve_clause -> calc_agent [bind_tools]
-> [conditional: tool_calls?] -> ToolNode -> calc_agent ...
-> retrieve_trend -> decide [conditional]
if BREACH:
-> challenge -> reconcile [conditional]
if confidence < 0.6 and _loop_count < 1:
-> retrieve_clause (loop back)
else:
-> memo_agent [bind_tools] -> ToolNode -> memo_agent -> END
else:
-> memo_agent [bind_tools] -> ToolNode -> memo_agent -> END
ReAct Tool Loop
calc_agentandmemo_agentare LLM nodes that see the full conversation history + available tools (CalculateFinancialRatio,DraftMemo).- If the LLM returns
tool_calls, the graph routes toToolNode, which invokes the correct tool implementation fromtools.py. ToolNodeoutput is appended to the message history and fed back to the same LLM node.- The loop ends when the LLM returns no more
tool_calls, and execution proceeds to the next business-logic node. - This pattern is not two separate one-off tool calls — it is genuine multi-turn reasoning with tool use at multiple points in the graph.
Conventions
- backend/nodes.py — All node implementations. Pure functions returning partial state updates.
- backend/prompts.py — All system prompts. Prompts are constants or factory functions, never inline strings in
nodes.py. - backend/graph.py — StateGraph builder with conditional edge routers.
- backend/tools.py — Pydantic tool schemas + deterministic fallback implementations.
- backend/state.py — TypedDict schema.
- backend/main.py — FastAPI app. Endpoints:
POST /upload(PDF) +POST /run(analysis). - backend/pdf_parser.py —
PyPDF2text extraction + LLM covenant extraction with validation. - backend/data/store.py — Document store. Reads from
covenants.jsonfor pre-loaded docs; stores uploaded docs in memory/JSON. - backend/data/covenants.json — Pre-fetched real SEC credit agreement covenant sections.
- frontend/app/page.tsx — Single Next.js page. Two input tabs: "Select Company" and "Upload Document".
- frontend/components/StepTrace.tsx — PRIMARY FEATURE. Visual graph trace showing every executed node in real-time.
- frontend/components/MemoDisplay.tsx — Formatted memo with decision badge and confidence bar.
- frontend/components/AnalysisForm.tsx — Dual-mode input form (dropdown + file upload).
API Contracts
POST /upload
Content-Type: multipart/form-data
file: <raw.pdf>
Response: { "company_id": str, "covenants": dict } — extracted data. On failure: HTTP 400.
POST /run
{"company_id": "MosaicCo", "covenant_type": "fixed_charge_coverage"}
Response: Full AgentState JSON. Frontend renders final_decision, final_confidence, memo, clause_citation, and challenge_arguments.
Tool Calling Requirement
The agent runs a ReAct-style tool loop using a shared ToolNode:
from langgraph.prebuilt import ToolNode
tools = [CalculateFinancialRatio, DraftMemo]
tool_node = ToolNode(tools)
Bound tools:
CalculateFinancialRatio— computes ratio from financial inputsDraftMemo— generates the final structured memo with citations
How it works:
calc_agentandmemo_agentare LLM nodes withvultr_llm.bind_tools(tools).- If the LLM emits
tool_calls, execution routes toToolNode, which executes the tool and returns the result. - The result is appended to the message history and fed back to the LLM node.
- When no more
tool_callsare present, the graph proceeds to the next node.
Rubric value: This demonstrates multiple distinct tool decisions (calculation + document generation) in a single run. The LLM actively chooses which tool to call and when. If the API returns no tool_calls, deterministic fallbacks in tools.py still return correct values, but a warning is logged.
Adversarial Layer Rules
challengenode ALWAYS activates onpreliminary_decision == "BREACH".reconcileproducesfinal_decisionandfinal_confidence.- If
final_confidence < 0.6and_loop_count < 1, route back toretrieve_clauseonce, then forced memo. - Challenge prompt must reference specific clause text (e.g., "Section 6.1 excludes non-recurring items").
- Generic challenges are disqualifying. The counter-argument must cite the exact clause and propose a specific adjustment.
Data Strategy
- Pre-fetch real SEC credit agreement exhibits into
covenants.json. - The
data/store.pymodule abstracts retrieval so the graph node always says "retrieving from document store" regardless of whether it's a pre-loaded JSON entry or an uploaded PDF extraction. - Uploaded PDFs are parsed, extracted by LLM, validated (threshold must be numeric > 0, formula must contain expected keywords), then stored temporarily.
Key Files
| File | Purpose |
|---|---|
backend/main.py |
FastAPI server, /upload + /run |
backend/graph.py |
LangGraph builder + edge routers |
backend/state.py |
TypedDict schema |
backend/nodes.py |
All graph nodes |
backend/prompts.py |
LLM prompts |
backend/tools.py |
Pydantic tool schemas + fallback math |
backend/pdf_parser.py |
PDF text extraction + LLM covenant parser |
backend/data/store.py |
Document store (pre-loaded + uploaded) |
backend/data/covenants.json |
Real covenant clauses (pre-fetched) |
backend/test_graph.py |
End-to-end graph tests |
frontend/app/page.tsx |
Main UI page |
frontend/components/StepTrace.tsx |
Visual graph trace — PRIMARY |
frontend/components/MemoDisplay.tsx |
Memo with badge + confidence |
frontend/components/AnalysisForm.tsx |
Dual-mode input form |
Environment
VULTR_SERVERLESS_INFERENCE_API_KEY— Required. From console.vultr.com > Serverless > Inference.LANGCHAIN_API_KEY— Optional but recommended. Enables LangSmith tracing of graph executions. Invaluable for demonstrating multi-step reasoning to judges.GRADIUM_API_KEY— Optional. For TTS "Play Memo" feature. Get 45,000 credits at gradium.ai + 100,000 with couponRAISE-2026.- Python 3.10+ (for
|union types,TypedDict, etc.) - Node.js 20+ (for Next.js frontend)
Judging Criteria Mapping
| Rubric Item | Weight | How Covenant Sentinel Satisfies It |
|---|---|---|
| Demo | 50% | Working software: multi-step agent with visible node trace, tool-calling integration, adversarial self-check, cited memo, and real PDF processing. The 1-minute video shows the entire flow. |
| Impact | 25% | Finance covenant monitoring is a real enterprise problem. Banks and credit officers deal with false-positive fatigue. |
| Creativity | 15% | Self-challenge/reconcile layer — an agent that argues with itself using its own retrieved clause text. |
| Pitch | 10% | 3 minutes on stage: run the demo scenario live, show StepTrace, read the memo. No slides. |
60-Second Demo Video Script
0:00–0:08: User opens app. Selects "Select Company" tab. Chooses "MosaicCo", "Fixed Charge Coverage". Clicks Run Analysis. Animated StepTrace: plan → retrieve_clause → calc_ratio → retrieve_trend → decide.
0:08–0:18: "Calculated Ratio: 1.48" (below 1.50). StepTrace decide in red. "Preliminary: BREACH (confidence: 0.72)".
0:18–0:35: Money shot. StepTrace → challenge. Panel: "Devil's Advocate Report" — "Section 6.1 excludes non-recurring items. The $50K restructuring charge in Q2 is likely one-off and should be excluded."
0:35–0:48: StepTrace → reconcile. "Final Decision: WATCH (confidence: 0.68)". Confidence bar drops.
0:48–0:58: Final memo with inline citation to Section 6.1.
0:58–1:00: Close card: "Covenant Sentinel — built on LangGraph + Vultr Serverless Inference in 24 hours."
Demo Scenario (Test Case)
- Company: MosaicCo
- Covenant: fixed_charge_coverage
- Inputs: EBITDA = 1.2M, Fixed Charges = 810K
- Ratio: 1.48
- Threshold: 1.50
- Preliminary: BREACH (confidence 0.72)
- Challenge: cites Section 6.1 exclusion, questions Q2 restructuring charge
- Reconcile: shifts to WATCH, confidence drops to 0.68
- Final memo: cites both raw calculation and challenge reasoning
Failure Modes (Handled)
- Tool call fails / no tool_calls → Deterministic fallback, log warning.
- Missing company/covenant → FastAPI 400.
- LLM refuses structured output → Set confidence to 0.5, route to memo.
- Adversarial infinite loop →
_loop_counthard-cap at 1. - Vultr API rate limit → Exponential backoff (1s, 2s, 4s), then 503.
- PDF upload fails / unreadable →
/uploadreturns 400: "Could not extract text from PDF." - LLM hallucinates threshold / formula → Validation layer rejects non-numeric thresholds or formulas missing "EBITDA"/"Debt".
Future Tool Pool (Post-Hackathon)
The current tool pool is intentionally scoped for 24-hour stability:
CalculateFinancialRatio— computationDraftMemo— document generation
Potential additions (not implemented for hackathon):
| Tool | Use Case |
|---|---|
PythonExecution |
General-purpose calculation fallback. Lightweight — takes a code string, runs it in a sandboxed interpreter. |
CalculateConfidence |
Explicit confidence scoring from data quality + trend direction. |
ValidateInputs |
Pre-calculation sanity check (positive EBITDA, non-zero denominators). |
WebSearch |
Fetch live market data, interest rates, or recent SEC filings. High demo risk due to network dependency. |
EDGARFetch |
Live SEC filing retrieval. Blocked by bot detection; requires Selenium or proxy. |
CheckClauseExclusions |
Parse clause text for specific exclusion keywords (extraordinary, non-recurring, one-time). |
MCPClient |
Model Context Protocol integration for enterprise systems (databases, Slack, ERP). Heavy infrastructure. |
Rationale for deferral: The rubric evaluates "calls tools" — not "calls many tools." Two genuine tool decisions (calculation + memo) satisfy the requirement. External tools add infrastructure overhead with minimal incremental judging value in a 60-second demo. They are documented here for post-hackathon expansion.
Code Style
- Python:
black,ruff, strictmypytyping. No baredictwhereTypedDictor Pydantic applies. - TypeScript:
strictmode, functional components, noany. - Never commit API keys.
.envis in.gitignore.
