Imported from Yaso2Go/tbhp (
.claude/rules/AGENTS.md). Install upstream withnpx skills add Yaso2Go/tbhp --skill rules. Copyright stays with the author.
AGENTS.md - Guidelines for AI Agent Development
This document provides guidelines for AI agents working on the Brain-Inspired Memory System project.
Project Goal
To create a terminal-based chat application that simulates a human-like memory system. The system should be able to:
- Determine if user input constitutes a "memory."
- Store memories with semantic structure (vector embeddings, tags, importance).
- Automatically link new memories to existing related memories in a graph structure.
- Manually trigger consolidation of memories (e.g., summarize clusters, prune unimportant ones) via a command.
- Retrieve relevant memories during conversation to provide context for LLM responses.
- Allow visualization of the memory graph.
Key Technologies and Libraries
- Local LLM: Recommended setup uses Ollama with
llama3.1:8b-instruct-q6_Kfor all tasks (conversation, classification, linking, summarization, pruning). One model, no extra downloads.- Alternative: Hugging Face
transformers(Mistral, Llama) with 4-bit quantization for VRAM efficiency on CUDA.
- Alternative: Hugging Face
- Embeddings: Sentence Transformers. Recommended:
intfloat/multilingual-e5-smallfor Arabic+English;BAAI/bge-small-en-v1.5for English-only. - Graph Database:
networkxfor in-memory graph representation and operations (with Neo4j as the persistent backend). Output to DOT format for external visualization. - Vector Storage: Simple Python lists and direct cosine similarity (from
scikit-learn), with embeddings stored in Neo4j. - Scheduler: Consolidation is triggered manually via a user command (APScheduler removed).
- Configuration:
python-dotenvfor managing environment-specific settings (like Neo4j credentials). Local LLM and Neo4j settings are primarily insrc/config.py.
Development Process
- Follow the Plan: Adhere to the established plan.
- Local LLM First: Core logic now uses the local Mistral model. Fallbacks to heuristics are in place if the local LLM fails to load or respond.
- Testing: Write unit tests for deterministic components and heuristics. Manually test the chat interface and LLM interactions regularly, especially after changes to prompts or LLM utility functions.
- Modularity: Maintain clean structure with components in respective files.
- Visualization: Memory graph export to DOT format is available via the
/visualizecommand.
LLM Prompts and Interactions
When using Ollama (recommended), the provider handles prompt formatting. For Hugging Face models, llm_utils.py uses chat templates. Instruct models typically use [INST] ... [/INST]-style prompts.
- Input Classification (Stage 1):
- System Message:
"You are an expert text classifier. Your task is to accurately classify user input into one of the predefined categories: memory, question, command, opinion, misc. You must output ONLY the label and a confidence score (a float between 0.0 and 1.0), separated by a comma. For example: 'memory, 0.9' or 'question, 0.95'. Pay close attention to the nuances of the input to make the best classification." - User Prompt Segment (simplified, actual prompt includes few-shot examples embedded):
Classify the following text as one of: memory, question, command, opinion, misc. Output ONLY the label and confidence (0-1), comma-separated. Here are some examples: - Input: 'I went to the park yesterday and saw a dog.' -> Output: memory, 0.9 - Input: 'What is the weather like today?' -> Output: question, 0.98 - Input: 'Tell me a joke.' -> Output: command, 0.92 - Input: 'I think pizza is better than pasta.' -> Output: opinion, 0.85 - Input: 'That is a blue car.' -> Output: misc, 0.7 (classified as misc because it's a simple statement of fact without clear memory, question, command, or opinion indicators) Input: '<user_input>' -> Output: - Expected LLM Output:
label, confidence_score(e.g.,memory, 0.85)
- System Message:
- Tag Extraction (Stage 2):
- System Message:
"You are an expert tag extractor. Output ONLY a comma-separated list of tags. For example: travel, diving, Dahab" - User Prompt Segment:
"Extract up to 5 relevant tags or topics from the following text. Output as a comma-separated list. Keep tags concise (1-3 words max per tag).\n\nInput: '<user_input>'" - Expected LLM Output:
tag1, tag2, tag3
- System Message:
- Importance Estimation (Stage 2):
- System Message:
"You are an expert significance assessor. Output ONLY a numerical score between 0.0 and 1.0. For example: 0.75" - User Prompt Segment:
"Rate the importance of the following memory content on a scale of 0.0 to 1.0 (0.0=trivial, 1.0=very important). Consider emotional impact, significance for future recall, and uniqueness. Output ONLY the numerical score.\n\nInput: '<user_input>'" - Expected LLM Output:
importance_score(e.g.,0.7)
- System Message:
- LLM Relation Score (Stage 3):
- System Message:
"You are an expert in contextual analysis. Output ONLY 'yes' or 'no', a comma, and a confidence score (e.g., yes, 0.85 or no, 0.7)." - User Prompt Segment:
"Are the following two memories logically related or describe similar contexts? Answer with ONLY 'yes' or 'no', followed by a comma and a confidence score (0.0-1.0) for your assessment. Example: yes, 0.9\n\nMemory 1: '<memory1_content>'\nMemory 2: '<memory2_content>'" - Expected LLM Output:
yes, 0.9orno, 0.75
- System Message:
- Summarization (Stage 4 - Consolidation - Placeholder):
- System Message:
"You are an expert at synthesizing information into concise summaries." - User Prompt Segment:
"Summarize the key themes and information from the following related memories into a single concise statement (1-2 sentences):\n\n- <memory_content_1>\n- <memory_content_2>\n- ..."
- System Message:
- Conversational Response (Stage 5 - Retrieval):
- System Message:
"You are a helpful AI assistant with access to a memory system. Use the provided memories to inform your response. If no relevant memories are found, answer the question generally." - User Prompt Segment (constructed in
main.py):Relevant past memories: - Memory <id1> (Importance: <imp1>, Timestamp: <ts1>): <retrieved_memory1_content> - Memory <id2> (Importance: <imp2>, Timestamp: <ts2>): <retrieved_memory2_content> ... Current user question: "<user_query>" Based on these memories (if any) and the current question, provide a thoughtful and context-aware response:
- System Message:
Important Notes for LLM:
- Ollama (recommended): Run
ollama run llama3.1:8b-instruct-q6_Kto ensure the model is available. No Hugging Face downloads. - Transformers: Ensure the model is downloaded/cached. 4-bit quantization requires
bitsandbytes. CUDA recommended (RTX 3080 Ti 12GB or better). - For pruning and linking: use structured JSON output (
decision,confidence,reason,risk) and policies likerisk != loworconfidence < 0.75→ don't prune automatically.
Coding Conventions
- Follow PEP 8 guidelines.
- Use type hinting.
- Write clear and concise comments where necessary.
Future Considerations (Out of Scope for Initial Build but good to keep in mind)
- Persistent storage for memory nodes and graph (e.g., SQLite, Neo4j).
- More sophisticated vector database (FAISS, Qdrant).
- UI for visualization instead of just DOT export.
- More advanced consolidation strategies.
- User authentication or multiple user profiles.
Remember to use the set_plan tool if the overall strategy needs to change and plan_step_complete after each step.
If you need to ask for user input, use request_user_input.
If you make significant changes to an approved plan, inform the user with message_user.
Memory Consolidation Agent (Updated)
- During consolidation, orphan and low-linked memories are considered for new links.
- For each candidate pair:
- Only pairs with high semantic similarity (cosine similarity ≥ 0.35) are considered.
- An LLM is asked if the memories should be linked, what the relation is, and to justify the link.
- A link is only created if the LLM gives a clear, logical justification (not generic or weak).
- This results in a more human-like, logical memory graph with only meaningful connections.
- Phase 2: Consolidation now also applies emotional decay (Fading Affect Bias) to all stored memories.
Phase 2: Emotional Memory & Neuromodulation System
Overview
Phase 2 adds an emotional processing layer that modulates every stage of the memory pipeline. The system mirrors neuroscience research on how emotion shapes memory formation, storage, and retrieval.
Pipeline Flow
User Input -> Sensory Register -> Emotion Detection (GoEmotions + VADER)
-> Amygdala (tagging + encoding boost + flashbulb check)
-> Neuromodulator Update (DA, NE, CORT, 5HT)
-> Working Memory (decay modulated by neuromodulators)
-> Attention Gate (emotion scorer from amygdala)
-> LTM (Neo4j, with emotional fields persisted)
-> Retrieval (mood-congruent similarity blending)
-> Consolidation (emotional decay with Fading Affect Bias)
Key Components
| Component | File | Purpose |
|---|---|---|
| EmotionDetector | src/memory/emotion_detector.py |
Hybrid emotion detection (GoEmotions + VADER + LLM fallback) |
| AmygdalaModule | src/memory/amygdala.py |
Emotional tagging, Yerkes-Dodson encoding boost, flashbulb detection |
| NeuromodulatorSystem | src/memory/neuromodulator.py |
4-channel neurotransmitter simulation (DA, NE, CORT, 5HT) |
| EmotionalDecayProcessor | src/memory/emotional_decay.py |
Fading Affect Bias, trauma resistance during consolidation |
| EmotionalMemoryConfig | src/core/config.py |
Pydantic configuration for all Phase 2 parameters |
MemoryNode Emotional Fields (Phase 2)
emotional_valence(float, -1.0 to +1.0): PAD Pleasure dimensionemotional_arousal(float, 0.0 to 1.0): PAD Arousal dimensionemotional_dominance(float, 0.0 to 1.0): PAD Dominance dimensionemotional_categories(List[str]): Detected categorical emotionsemotional_intensity(float, 0.0 to 1.0): Overall emotional intensityis_flashbulb(bool): Whether this is a flashbulb memoryencoding_boost(float, 1.0 to 2.0): Encoding strength multiplier
New Chat Commands (Phase 2)
/mood- Display current neuromodulator levels and system state/emotion- Display last detected emotion and amygdala statistics
Configuration
All Phase 2 settings use the EMOTIONAL__ env prefix (see .env.example). Key settings:
EMOTIONAL__EMOTION_MODEL_ID- GoEmotions model (default:SamLowe/roberta-base-go_emotions)EMOTIONAL__FLASHBULB_AROUSAL_THRESHOLD- Flashbulb trigger threshold (default: 0.8)EMOTIONAL__NEGATIVE_DECAY_MULTIPLIER- FAB multiplier (default: 1.5x faster for negative)EMOTIONAL__HOMEOSTASIS_TURN_THRESHOLD- Homeostasis correction after N turns (default: 10)
