The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Hipocampo listing page.
Persistent memory for autonomous AI agents · PostgreSQL 17 + pgvector · Hybrid Search · MCP Server
⚠️ Transport Note: SSE transport is deprecated since MCP spec 2025-03-26. Hipocampo now uses Streamable HTTP (single endpoint
/mcp) as the recommended remote transport. SSE (/sse) remains available for backward compatibility but will be removed in a future release.
Hipocampo runs as a free MCP server on Hugging Face Spaces. Connect from any MCP client:
🧪 Interactive Playground: Try saving and searching memories from your browser at https://alexbell1-hipocampo-mcp.hf.space/ — no registration or MCP client needed.
⚠️ Important: The Hugging Face free tier is ephemeral — data is lost on restart/deploy. This instance is intended for testing only. For persistent storage, run Hipocampo locally (see Quick Start) or connect an external database (Neon, Supabase, etc.).
Embedding model: sentence-transformers/all-MiniLM-L6-v2 (384 dims) via Hugging Face Inference API (free, no credit card required).
Una sola línea. La terminal hace todo: PostgreSQL + pgvector, embeddings, base de datos, venv, clientes MCP, servicio systemd con timer de mantenimiento automático.
🪟 ¿Usas Windows?
install.shes un script de Linux/macOS. En Windows necesitas WSL2 (Windows Subsystem for Linux):Abre la terminal de WSL (Ubuntu), actualiza los paquetes y vuelve a ejecutar el instalador:
⚠️ En Windows puro (CMD/PowerShell) el instalador NO funciona. Sin WSL2 verás errores como
Package 'python3-venv' has no installation candidateoapt: command not found.
Máquinas sin interacción (VPS, contenedores):
| Fase | Qué instala | 🕐 |
|---|---|---|
| ① | Diagnóstico: OS, gestor de paquetes, RAM, disco | ~2s |
| ② | PostgreSQL 17 + pgvector (apt · dnf · pacman · brew) | ~15s |
| ③ | Base de datos + esquema: 10 tablas, HNSW + GIN, ownership | ~3s |
| ④ | Embeddings: Ollama local (qwen3-embedding:0.6b) o API externa | ~60s |
| ⑤ | Python .venv + pip + archivo .env | ~10s |
| ⑥ | Clientes MCP: OpenCode · Claude · Gemini/Antigravity · Cursor · VS Code · Windsurf | ~2s |
| ⑦ | Servicio systemd + timer semanal de mantenimiento automático | ~1s |
| ⑧ | Autodiagnóstico: health · save · search · cleanup | ~3s |
| ✅ Idempotente | Vuelve a ejecutarlo sin miedo — repara ownership, actualiza repo y configs |
| ✅ 6 clientes MCP | OpenCode, Claude, Gemini/Antigravity, Cursor, VS Code, Windsurf |
| ✅ Mantenimiento automático | Timer semanal (domingo 03:00) con Persistent=true — catch-up si la PC estaba apagada |
| ✅ Sin root | Todo en ~/.local/share/hipocampo |
| ✅ Opciones | --unattended, --embed-api, --install-dir, --db-user, --no-clients, --no-timer, --no-ollama |
| ✅ Desinstalación limpia | bash uninstall.sh — para servicios, BD opcional, clientes, archivos |
| ✅ Compilación desde fuente | Fallback si el paquete pgvector no está en el repositorio |
Hipocampo is an advanced dual-memory persistence architecture designed for autonomous AI agents. By maintaining both technical knowledge and user profiling data across sessions, Hipocampo provides a reliable, stateful context that enables agents to learn, adapt, and scale efficiently.
Built on top of PostgreSQL 17 with pgvector, it features BIRE v3.7 — a hybrid retrieval engine combining semantic embeddings (1024d), lexical expansion, and GIN trigram search with dynamic score fusion. Also includes Sparse Selective Caching (SSC) as an experimental pipeline.
Hipocampo already reduces context through SSC (selective retrieval). But even the top-5 most relevant memories can consume 500-2000+ tokens when concatenated — a significant portion of any LLM's context window.
Hybrid compression adds a second reduction layer:
Real impact: If you call compress_hipocampo before every search_hipocampo → LLM round-trip, you save 200-800 tokens per interaction. At scale (hundreds of queries), this translates to meaningful cost reduction and faster responses.
memoria_vectorial) and user profile data (memory_items), each utilizing 1024-dimensional embeddings.qwen3-embedding:0.6b by default), query expansion, GIN trigram, and composite scoring — used by all MCP tools.compress_hipocampo MCP tool.link_hipocampo, graph_hipocampo, path_hipocampo MCP tools.trigger:php, trigger:chartjs, trigger:tomcat) — when the agent starts working in that context, it searches for matching automatica rules and reactivates past errors before making the same mistake. This mirrors the biological hippocampus: a partial cue (project + language) triggers full memory retrieval of the error and its solution. Automatic rules are permanent — never compressed, never deleted. set_nivel_hipocampo(id, nivel) + consolidate_hipocampo tools included.automatica rule capturing the exact cause, symptom, and fix. Uses immune economy: pre-change snapshots are cheap episodica (auto-compressed if no damage), post-break rules are permanent automatica. Pre-loaded with fragile file catalog — header.php, conexion.php, utils.php, auth.php, etc. Agents search trigger:regression trigger:<file> before every edit to learn what other agents broke before.search_code(query, language) — returns real code snippets with file paths and line numbers, not just summaries.final_score = relevance × exp(-λ × days) with λ=0.05 configurable and 20% floor. Recent knowledge naturally outranks old memories.diversity_lambda × relevance - (1-diversity_lambda) × max_similarity_to_selected. Configurable in hipocampo_hybrid_config.json.decay_hipocampo now archives old episodica memories to memoria_historica (cold storage) when they exceed age thresholds. Protected levels: automatica, semantica, critico — never archived. New critico parameter on save_hipocampo for mission-critical memories. restaurar_historica(id) restores cold memories back to active tier.memory_access table tracks per-record access frequency. BIRE search applies a fatigue boost: boost = min(15, 5·log1p(accesses_7d))·e^(-age_hours/168). Frequently accessed memories naturally rank higher — mimicking how the human brain strengthens neural pathways through repeated recall.hipocampo_budget(dry_run) manages three storage tiers — HOT (embedding present, full semantic search), WARM (embedding=NULL, text-only search), COLD (memoria_historica archive). Hot tier cap: 5000 records. When the cap is exceeded, oldest episodica memories are automatically demoted to WARM. restaurar_historica(id) restores any cold memory back to active.save_hipocampo runs _detectar_contradicciones() using negation-probe embeddings to detect factual contradictions with existing memories. When detected: logs a warning and creates a contradicts link — never blocks the save. contradicciones_hipocampo(id) performs on-demand contradiction audits across the memory graph.hipocampo_watch.py watches configured directories for file changes and auto-reindexes modified files via index_project. Managed by hipocampo-watch.timer (10-minute interval). MCP tools: list_watch_dirs, add_watch_dir(path, patterns), remove_watch_dir(path), reindex_now(path?).graph_hipocampo() and path_hipocampo() auto-reinforce traversed links. New decay_hipocampo(dry_run) tool for graph maintenance. Columns: last_accessed, reinforced_at.trade_knowledge (nunca se decae), clasificación automática de reusabilidad (high→promoción a semántica), perfiles de decaimiento por dominio (infra 180d, proyecto 90d, temporal 14d), y recordatorio trimestral con review_trade_knowledge(). 266 memorias clasificadas en migración automática.preload_context(project_path) extracts meaningful keywords from the project path, searches relevant memories, and returns a compressed summary — ideal for session start.compress_hipocampo auto-estimates token budget and adjusts k dynamically. budget_ratio parameter gives fine-grained control over output size.save_hipocampo(..., auto_link=True) auto-discovers semantically similar memories (>0.75 cosine) and creates similar edges in the memory graph.hipocampo_health() checks the HNSW index on startup and auto-creates it if missing — no more manual CREATE INDEX commands.You might wonder why Hipocampo uses PostgreSQL 17 with pgvector instead of a lighter stack like SQLite. The answer: hybrid search requires more than vector similarity alone.
Hipocampo's retrieval pipeline combines pgvector (HNSW) for semantic search, pg_trgm (GIN) for lexical expansion, and ILIKE for fallback — fused into a single weighted score. SQLite extensions like sqlite-vec offer vector search, but lack:
With ~1,100+ records across two memory tables and growing, Hipocampo needs a database that scales without sacrificing retrieval quality. PostgreSQL + pgvector isn't "heavy" for the sake of it — it's the minimum viable stack to deliver the hybrid accuracy that BIRE and SSC require.
Hipocampo enables AI agents to learn from mistakes across sessions using a simple cycle:
Real example: An agent tries flatpak install npm and fails. It saves the error to Hipocampo: "npm is a Node.js package manager, not a Flatpak package. Use npm directly." Next time the same command is attempted, the agent finds this record and knows the solution immediately — without repeating the mistake.
Over time, the agent's error knowledge base grows organically. Each failure makes future sessions smarter. This turns Hipocampo from a simple archive into a continuous learning system for AI agents.
Going beyond reactive learning, Hipocampo v4.1 introduces trigger-based automatic rules that fire before the agent writes a single line of code:
How to implement:
This mirrors the biological hippocampus: a partial cue triggers full memory retrieval — the brain doesn't wait for the error to happen before remembering it hurts.
Sometimes the agent breaks code that was working fine — not repeating an old error, but creating a new one. Hipocampo v4.2 implements a 3-step immune cycle that mirrors how the body generates antibodies:
Fragile file catalog (pre-loaded): header.php, conexion.php, utils.php, auth.php, db_connection.php — these files have cascading dependencies. One wrong edit breaks dozens of pages. Hipocampo ships with fragile file rules so agents know what to handle with care.
Before every edit: search_hipocampo("trigger:regression trigger:<file> trigger:<project>") — learn what other agents broke on this file before you touch it.
To enable this behavior, you need to instruct your agent to use the cycle above. This is done by adding instructions to the agent's configuration file, depending on the client:
| Agent | Configuration file | Example |
|---|---|---|
| OpenCode | AGENTS.md (project root) or ~/.opencode/AGENTS.md | See example |
| Claude Code | CLAUDE.md or ~/.claude/CLAUDE.md | Similar approach |
| Cursor | .cursorrules | Add instructions in plain text |
| Windsurf | .windsurfrules | Same structure |
| Cline | CLINE.md | Same structure |
Minimal example for AGENTS.md / CLAUDE.md:
💡 Tip: For MCP-native agents (OpenCode, Claude Code), Hipocampo tools are available directly. For others, use the HTTP endpoint or CLI scripts.
BIRE search results can be large (especially after increasing the per-result display limit to 8000 characters). OpenCode clients truncate tool output by default at 2000 lines / 51 KB. You'll see ...N bytes truncated... at the bottom when this happens. The full output is saved to a file for later reading, but to avoid truncation entirely, add to your opencode.json or ~/.config/opencode/opencode.jsonc:
Restart OpenCode for the change to take effect.
📄 Want your agent to use ALL of Hipocampo? Append AGENTS_TEMPLATE.md sections into your AGENTS.md / CLAUDE.md / .cursorrules — safe to share, no credentials.
💡 ¡Recomendado! En vez de seguir los pasos manuales, ejecuta el instalador automático:
El instalador configura PostgreSQL + pgvector, embeddings, base de datos, venv, clientes MCP y el timer de mantenimiento. Ver Auto-Installer v6.0 para más detalles.
pgvector and pg_trgm extensions enabled)Hipocampo provides specialized scripts to interact with the core engine:
The core of Hipocampo is backed by a relational and vector hybrid design:
BIRE (Búsqueda Integrada por Relevancia Expansiva) is the default search engine used by all MCP tools. It combines vector and lexical search with dynamic score fusion:
An SSC (Sparse Selective Caching) pipeline is also available as an experimental alternative:
Hipocampo includes a fully functional FastMCP server, allowing LLM agents to autonomously read and write memories.
Memory Operations:
search_hipocampo(query, session_id?): Unified semantic and lexical search (auto-records metrics). Optionally filter by session.quick_hipocampo_search(query): Shorthand alias for rapid queries.preload_context(project_path, k=8): Extract keywords from project path, search relevant memories, return compressed summary. Ideal for session initialization.compress_hipocampo(query, k=5, method="hybrid", budget_ratio=1.0, include_metadata=False): Search + hybrid compression with context budget awareness. Auto-estimates tokens and adjusts k dynamically. Three methods: "hybrid" (recommended), "extractive" (fastest, no API cost), "llm" (highest quality).save_hipocampo(content, memory_type, code, categories, session_id?, force?, auto_link=False, nivel="episodica"): Persist data into memoria_vectorial. Supports session isolation, auto-dedup, auto-linking, and hierarchical memory levels.profile_hipocampo(summary, extra, categories): Store personal or event-driven user data (memory_items).save_hipocampo now supports critico=True parameter to protect mission-critical memories from decay and archiving.Memory Graph (v4.0):
link_hipocampo(source_id, target_id, relation_type, weight): Create a directed edge between two memories. Relation types: related, follow_up, part_of, references, similar, chain.unlink_hipocampo(id / source+target+type): Remove edge(s) from the memory graph.graph_hipocampo(node_id, depth=2): BFS tree traversal from a root node. Use node_id=0 for an overview of all connected nodes and edge counts.path_hipocampo(from_id, to_id, max_depth=5): Find the shortest BFS path between two memories.Code RAG (v4.0):
index_project(project_path, force=False): Scan and index source code files as semantic embeddings. Incremental — only re-indexes changed files (by mtime). Supports PHP, JS, TS, Python, SQL, HTML, CSS, JSON, YAML.search_code(query, k=5, language=""): Vector search specifically in indexed code snippets. Returns real code with file paths, language, and line numbers.CRUD Operations:
update_hipocampo(id, content?, memory_type?, code?, categories?): Update an existing memory. Regenerates embedding if content changes.delete_hipocampo(id): Permanently delete a memory by ID.set_nivel_hipocampo(id, nivel): Promote/demote a memory between hierarchical levels (episodica, semantica, automatica).consolidate_hipocampo(min_age_days=7, dry_run=True): Migrate old episodic memories to semantic level with optional content compression.Self-Diagnosis & Auto-Repair:
hipocampo_health(): Full system health check (PostgreSQL, embedding API, disk, extensions, HNSW index).hipocampo_auto_repair(): Automatically repairs detected issues (restart PostgreSQL, create missing tables, create HNSW index).Performance Optimization (Fase 2):
hipocampo_stats(): Query performance metrics, latency analysis, and optimization recommendations.hipocampo_tune(): Auto-adjusts BIRE/SSC thresholds and hybrid weights based on real usage data.Memory Maintenance (Fase 3):
hipocampo_dedup(merge): Detects and merges duplicate memories (exact + semantic via cosine similarity).hipocampo_checkpoint(dry_run): Logarithmic checkpointing to compress old memories.hipocampo_maintenance(): Full maintenance cycle (repair → dedup → checkpoint → tune).Time Decay:
Active Forgetting & Tiering (v5.0):
decay_hipocampo(dry_run=True): Extended to archive old episodica memories to memoria_historica (cold storage). Protected: automatica, semantica, critico. Dry run shows what would be archived.hipocampo_budget(dry_run=True): Shows memory distribution across HOT/WARM/COLD tiers. Hot cap: 5000. When exceeded, oldest episodica are auto-demoted.restaurar_historica(id): Restore a cold memory from memoria_historica back to active memoria_vectorial.contradicciones_hipocampo(id=None): On-demand contradiction audit. With ID: checks one memory. Without: scans all memories for contradictions.Trade Knowledge Preservation (v4.3):
review_trade_knowledge(dry_run=True): Lists infrastructure memories approaching their decay limit (>150 days). Use to manually reinforce trade knowledge before it auto-decays.list_trade_knowledge(): Lists all memories tagged as trade_knowledge=true with their reusability and domain profile.File Watcher (v5.0):
list_watch_dirs(): List all directories being watched for auto-reindexing.add_watch_dir(path, patterns=["*.php","*.py","*.js"]): Add a directory to the watch list.remove_watch_dir(path): Remove a directory from the watch list.reindex_now(path=None): Trigger immediate reindex of watched files (or all if no path given).Webhook Watches:
watch_hipocampo(pattern, webhook_url): Register a webhook that fires on save/update/delete events matching a text pattern.unwatch_hipocampo(id): Remove a registered webhook.list_watches(): List all registered webhooks and their targets.For advanced configuration, please refer to the MCP Server Guide.
DB connection, config loading, and embedding generation are centralized in the hipocampo package:
All scripts in scripts/ import from hipocampo.db instead of duplicating the boilerplate. The MCP server also imports search/health/stats/dedup/checkpoint functions directly — no subprocess calls.
Before: Each MCP search spawned subprocess.run() → fork Python interpreter → re-import everything → connect DB → generate embedding → run query → parse stdout. That's ~200–500ms of process + serialization overhead alone.
After: Direct function call within the same process. The DB connection pool, OpenAI client, and modules are already cached. Overhead drops to microseconds.
For individual searches the difference is marginal (~200ms), but for hipocampo_maintenance() it previously ran 4 serial subprocess forks — now it's one direct call per phase, saving ~1–2 seconds.
The MCP server now runs all 16 tools as async Python coroutines in HTTP mode, and uses a PostgreSQL connection pool instead of creating a new connection per call:
Before:
connect() latency on every callsearch froze the server for all concurrent clientstoo many connections on the databaseAfter:
init_pool(minconn=1, maxconn=10) creates a ThreadedConnectionPool at server startup — connections are reused across calls, handshake happens onceasync def — blocking I/O (DB queries, embedding API) runs in asyncio.to_thread(), freeing the event loop for other requests_PooledConnection proxy transparently returns connections to the pool when .close() is called — zero caller-side changesImpact: Concurrent requests no longer block each other; PostgreSQL connection overhead drops from ~10–50ms per call to near zero.
Integration Tests:
@pytest.mark.integration) start the server in stdio mode and verify tools/list, resources/list, and a real search callBefore:
DB_HOST or embedding config → server started without errors, failed with cryptic fe_sendauth / 401 on the first queryexcept Exception: logger.error("msg: %s", e) — no traceback, impossible to tell if it was a DB, network, or validation failureAfter:
validate_config() runs at startup and logs clear warnings for each missing variable. init_pool() and get_conn() reject early with messages like "PostgreSQL connection incomplete: DB_HOST, DB_USER not configured in .env"embedding_limiter (30/min — shields embedding API cost), tool_limiter (60/min — shields PostgreSQL), watch_limiter (20/min). Clients get "⏳ Too many requests. Limit: 30 per 60s. Wait 12s."_tool_err() helper differentiates by exception type: psycopg2.Error → logger.exception() with full traceback, ValueError / TypeError → logger.warning() (client error), others → logger.exception(). _fire_webhooks catches urllib.error.URLError separatelyImpact: Failures are caught before they reach the database, costs are capped, and logs are actionable — you know instantly if it's a misconfiguration, a network blip, or a code bug.
Before:
get_embedding() failed on the first embedding API timeout or rate limit — no retry at allsys.argv parsing — no --help, no type validation, inconsistent interfacesAfter:
get_embedding() uses tenacity with wait_exponential(mult=1, min=1, max=30), 5 attempts, retrying only on RateLimitError/APITimeoutError/APIConnectionError/InternalServerError. No retry on AuthenticationError or BadRequestError. Each retry is logged at warning levelargparse with --help, typed arguments, and consistent names: hipocampo_mcp_server.py --http 8001.pre-commit-config.yaml with ruff lint+format (pre-commit) and pytest (pre-push). pyproject.toml configures ruff with line-length 120Impact: The server tolerates transient API failures without the client seeing errors. CLI is self-documenting. Every commit is verified before reaching GitHub — no more broken tests on main.
compress_hipocampo Tool, and symlink-based Structure (v3.9)Before:
scripts/ were independent copies of the repo — each git pull required manual sync, and new files like hipocampo_compress.py were missingscripts/ directory — no separation from repo fileshipocampo/ Python package was also a copy: load_config() looked for .env in project_root/.env instead of ~/.hipocampo/.env, loading incorrect credentials (alex/hipocampo123)compress operations failed with a generic exceptionquery_stats table existed in esquema.sql but was never auto-created — hipocampo_health reported DEGRADEDregister_vector(conn) on the _PooledConnection from get_conn() — psycopg2 rejected it with TypeError, breaking all vector operationsinitialize handshake, failing with Invalid request parametersNow:
~/.hipocampo/ is the canonical home: repo/ (git clone), ~/.hipocampo/scripts/ → repo/scripts/ (symlink), ~/.hipocampo/hipocampo/ → repo/hipocampo/ (symlink). Local user scripts moved to ~/.hipocampo/local_scripts/. git pull on repo/ auto-updates everything_find_env() in db.py loads .env in deterministic order: ENV_PATH env var → ~/.hipocampo/.env (explicit user config) → project_root/.env (Docker/Fly). No more wrong credentialsensure_stats_table() runs at module import time in the MCP server — query_stats table auto-created on starthipocampo_compress.py with explicit exception handling: RateLimitError, APITimeoutError, APIConnectionError, APIStatusError → immediate fallback to extractive compression. All other errors → fallback with distinct warning levelregister_vector(conn) removed from hipocampo_search.py, hipocampo_checkpoint.py, hipocampo_calibrate.py, mm_brain_tool.py — get_conn() already registers the vector adapter on the real connectionmcp[client] SDK (stdio_client + ClientSession + initialize()) — proper MCP 2025-03-26 handshakeImpact: Zero-touch maintenance after git pull. Transient embedding API errors degrade gracefully. Config loading is deterministic and secure. Vector operations work reliably. Tests follow the official MCP protocol.
Unattended consolidation, decay, and pruning — the memory system now cleans itself. Three complementary mechanisms:
The MCP server can run a maintenance loop every 24h (configurable). Off by default — activate with an environment variable:
| Env var | Default | Purpose |
|---|---|---|
HIPOCAMPO_AUTO_MAINTENANCE | false | Enables the scheduler in the HTTP lifespan |
HIPOCAMPO_AUTO_MAINT_INTERVAL_S | 86400 | Seconds between maintenance cycles |
HIPOCAMPO_MAINT_MIN_AGE_DAYS | 7 | Min age to consolidate episódica → semántica |
HIPOCAMPO_MAINT_DECAY_MIN_AGE_DAYS | 60 | Min age to archive unaccessed episódica (active forgetting) |
Each cycle runs: consolidation (episodic → semantic promotion), decay (link weight half-life 90d + active forgetting of unaccessed episodic memories), dedup merge, and access-log purge (>30d). All protections intact: automatica, semantica, critico, and linked memories are never archived.
Every 50 saves (configurable via HIPOCAMPO_SAVE_TRIGGER_EVERY, 0 disables), a background thread runs a lighter cycle: consolidation + decay + purge — no dedup merge (irreversible). The system cleans itself in proportion to how much it's used, no external services required.
Persistent=true (recommended for desktops)For machines that power off at night, cron loses scheduled runs. A user-level systemd timer with Persistent=true catches up on the missed run as soon as the PC boots:
Install the weekly timer (Sunday 03:00, catch-up on boot):
The CLI reuses _run_maintenance_cycle() from the MCP server — the exact same code path as the internal scheduler, zero logic duplication.
decay_hipocampo never ran: the memory-level query mixed timestamptz and text in a COALESCE (COALESCE(max(accessed_at), metadatos->>'date')) → PostgreSQL error types timestamp with time zone and text cannot be matched. Fixed by casting (NULLIF(metadatos->>'date',''))::timestamptz. Same bug fixed in hipocampo_budget.min_age_days was ignored: decay_hipocampo(dry_run=False, min_age_days=30) accepted the parameter but Part 2 (active forgetting) had no age filter in SQL — it would archive episodic memories of any age. Now the age threshold is parameterized in the query.If this project helps you, consider supporting its development:
Optimizations applied in July 2026 to address latency and threshold drift:
get_embedding() now uses an LRU cache (128 entries) — repeated queries for identical text skip the embedding API call entirely, saving ~450ms each. The OpenAI client is also reused across calls instead of being recreated.
SSC_TOP_K reduced from 20 → 15: fewer vector results per table means faster vector searchhnsw.ef_search = 20: lower HNSW breadth-of-search for approximate (faster) nearest neighbors (default was 40)CONFIANZA_ALTA lowered from 70 → 60: early exit from the SSC pipeline sooner when vector results are already goodregister_vector cached per connection to avoid redundant SQL introspectionalpha reset from 0.6 → 0.5 (balanced 50% vector + 50% lexical), vectorial_confidence_min from 0.75 → 0.70hipocampo_tune()) now capped: alpha stays within 0.4–0.6, confidence within 0.5–0.75register_vector() and detects the pgvector/PG17 indam incompatibility with a clear upgrade messageHipocampo includes 103 unit tests covering all core logic and MCP integration:
| Test file | What it covers |
|---|---|
tests/test_search.py | Query expansion (stem map + synonyms), score fusion with dynamic alpha, temporal decay (5%/week), result formatting |
tests/test_autotag.py | All 17 tag rules, 16 category rules, memory_type auto-detection |
tests/test_dedup.py | Cosine similarity (including 1024-dim vectors), exact and semantic duplicate detection logic |
tests/test_checkpoint.py | Age scale classification, project grouping, summary generation |
tests/test_mcp_integration.py | 6 schema tests (tool registration, annotations, params, async signature) + 3 live integration tests (stdio server, mcp[client] SDK) |
tests/test_rate_limit.py | Sliding-window rate limiter: acquire/release, prune, stats, default limiters |
tests/test_db.py | Config validation: missing DB_HOST, embedding API config, comprehensive coverage |
Tests run automatically on every push via GitHub Actions on Python 3.11–3.13.
This project is licensed under the MIT License.
Memoria persistente para agentes de IA autónomos · PostgreSQL 17 + pgvector · Búsqueda Híbrida · Servidor MCP
⚠️ Nota de Transporte: SSE está deprecado desde spec MCP 2025-03-26. Hipocampo ahora usa Streamable HTTP (endpoint único
/mcp) como transporte remoto recomendado.
Hipocampo corre como servidor MCP gratuito en Hugging Face Spaces. Conéctate desde cualquier cliente MCP:
🧪 Playground interactivo: Prueba guardar y buscar recuerdos desde el navegador en https://alexbell1-hipocampo-mcp.hf.space/ — sin registro ni cliente MCP.
⚠️ Importante: El tier gratuito de Hugging Face es efímero — los datos se pierden al reiniciar/desplegar. Esta instancia es solo para pruebas. Para persistencia real, ejecuta Hipocampo localmente o conecta una base externa.
Una sola línea. La terminal hace todo: PostgreSQL + pgvector, embeddings, base de datos, venv, clientes MCP, servicio systemd con timer de mantenimiento automático.
🪟 ¿Usas Windows?
install.shes un script de Linux/macOS. En Windows necesitas WSL2 (Windows Subsystem for Linux):Abre la terminal de WSL (Ubuntu), actualiza los paquetes y vuelve a ejecutar el instalador:
⚠️ En Windows puro (CMD/PowerShell) el instalador NO funciona. Sin WSL2 verás errores como
Package 'python3-venv' has no installation candidateoapt: command not found.
Máquinas sin interacción (VPS, contenedores):
| Fase | Qué instala | 🕐 |
|---|---|---|
| ① | Diagnóstico: OS, gestor de paquetes, RAM, disco | ~2s |
| ② | PostgreSQL 17 + pgvector (apt · dnf · pacman · brew) | ~15s |
| ③ | Base de datos + esquema: 10 tablas, HNSW + GIN, ownership | ~3s |
| ④ | Embeddings: Ollama local (qwen3-embedding:0.6b) o API externa | ~60s |
| ⑤ | Python .venv + pip + archivo .env | ~10s |
| ⑥ | Clientes MCP: OpenCode · Claude · Gemini/Antigravity · Cursor · VS Code · Windsurf | ~2s |
| ⑦ | Servicio systemd + timer semanal de mantenimiento automático | ~1s |
| ⑧ | Autodiagnóstico: health · save · search · cleanup | ~3s |
| ✅ Idempotente | Vuelve a ejecutarlo sin miedo — repara ownership, actualiza repo y configs |
| ✅ 6 clientes MCP | OpenCode, Claude, Gemini/Antigravity, Cursor, VS Code, Windsurf |
| ✅ Mantenimiento automático | Timer semanal (domingo 03:00) con Persistent=true — catch-up si la PC estaba apagada |
| ✅ Sin root | Todo en ~/.local/share/hipocampo |
| ✅ Opciones | --unattended, --embed-api, --install-dir, --db-user, --no-clients, --no-timer, --no-ollama |
| ✅ Desinstalación limpia | bash uninstall.sh — para servicios, BD opcional, clientes, archivos |
| ✅ Compilación desde fuente | Fallback si el paquete pgvector no está en el repositorio |
Hipocampo es una arquitectura avanzada de persistencia de memoria dual diseñada para agentes de Inteligencia Artificial. Al mantener tanto el conocimiento técnico como los datos del perfil del usuario entre sesiones, Hipocampo proporciona un contexto con estado confiable que permite a los agentes aprender, adaptarse y escalar eficientemente.
Construido sobre PostgreSQL 17 y pgvector, utiliza BIRE v3.7 — un motor híbrido que combina embeddings semánticos (1024d), expansión léxica y búsqueda GIN trigram con fusión dinámica de puntuación. Incluye también Caché Selectivo (CS/SSC) como pipeline experimental.
Hipocampo ya reduce el contexto mediante SSC (búsqueda selectiva). Pero incluso las 5 memorias más relevantes pueden consumir 500-2000+ tokens al concatenarse — una porción significativa de la ventana de contexto del LLM.
La compresión híbrida añade una segunda capa de reducción:
Impacto real: Si usas compress_hipocampo antes de cada search_hipocampo → LLM, ahorras 200-800 tokens por interacción. A escala (cientos de consultas), esto se traduce en reducción significativa de costos y respuestas más rápidas.
memoria_vectorial) y datos de perfil (memory_items), ambas utilizando embeddings de 1024 dimensiones.qwen3-embedding:0.6b por defecto), expansión de consulta, GIN trigram y puntuación compuesta — usado por todas las tools MCP.compress_hipocampo.link_hipocampo, graph_hipocampo, path_hipocampo.trigger:php, trigger:chartjs, trigger:tomcat) — cuando el agente comienza a trabajar en ese contexto, busca reglas automatica coincidentes y reactiva errores pasados antes de cometer el mismo error. Esto replica el hipocampo biológico: una pista parcial (proyecto + lenguaje) dispara la recuperación completa del error y su solución. Las reglas automáticas son permanentes — nunca se comprimen, nunca se eliminan. Tools: set_nivel_hipocampo(id, nivel) + consolidate_hipocampo.automatica permanente que capture la causa exacta, el síntoma y la solución. Usa economía inmune: los snapshots pre-cambio son episodica baratos (se autocomprimen si no hubo daño), las reglas post-rotura son automatica permanentes. Precargado con catálogo de archivos frágiles — header.php, conexion.php, utils.php, auth.php, etc. El agente busca trigger:regresion trigger:<archivo> antes de cada edición para aprender lo que otros agentes rompieron antes.search_code(consulta, lenguaje) — devuelve código real con ruta de archivo y números de línea.score_final = relevancia × exp(-λ × días) con λ=0.05 configurable y piso 20%. El conocimiento reciente pesa naturalmente más.diversity_lambda × relevancia - (1-diversity_lambda) × max_similitud_a_seleccionados. Configurable en hipocampo_hybrid_config.json.decay_hipocampo ahora archiva memorias episodica antiguas en memoria_historica (almacenamiento frío) al superar umbrales de edad. Niveles protegidos: automatica, semantica, critico — nunca se archivan. Nuevo parámetro critico en save_hipocampo para memorias críticas. restaurar_historica(id) restaura memorias frías al tier activo.memory_access rastrea frecuencia de acceso por registro. BIRE aplica un boost de fatiga: boost = min(15, 5·log1p(accesos_7d))·e^(-edad_horas/168). Las memorias accedidas con frecuencia en una ventana de 7 días suben naturalmente en el ranking — imitando cómo el cerebro fortalece vías neuronales mediante la recuperación repetida.hipocampo_budget(dry_run) gestiona tres tiers — HOT (embedding presente, búsqueda semántica completa), WARM (embedding=NULL, búsqueda solo por texto), COLD (memoria_historica archivo). Cap del tier hot: 5000 registros. restaurar_historica(id) restaura memorias frías al tier activo.save_hipocampo ejecuta _detectar_contradicciones() usando embeddings de sonda de negación para detectar contradicciones factuales. Cuando detecta: registra warning y crea enlace contradicts — nunca bloquea el guardado. contradicciones_hipocampo(id) realiza auditorías de contradicción bajo demanda.hipocampo_watch.py monitorea directorios configurados y auto-reindexa archivos modificados via index_project. Gestionado por hipocampo-watch.timer (intervalo 10 minutos). Tools MCP: list_watch_dirs, add_watch_dir(path, patterns), remove_watch_dir(path), reindex_now(path?).graph_hipocampo() y path_hipocampo() refuerzan automáticamente los enlaces atravesados. Nueva tool decay_hipocampo(dry_run) para mantenimiento del grafo. Columnas: last_accessed, reinforced_at.preload_context(ruta_proyecto) extrae keywords del proyecto, busca memorias relevantes y devuelve resumen comprimido. Ideal al inicio de sesión.compress_hipocampo auto-estima tokens y ajusta k dinámicamente. budget_ratio da control fino sobre el tamaño de salida.save_hipocampo(..., auto_link=True) descubre recuerdos semánticamente similares (>0.75 cosine) y crea aristas similar en el grafo.hipocampo_health() verifica el índice HNSW al arrancar y lo crea si falta — sin comandos CREATE INDEX manuales.Quizás te preguntes por qué Hipocampo usa PostgreSQL 17 con pgvector en lugar de algo más ligero como SQLite. La respuesta: la búsqueda híbrida necesita más que solo similitud vectorial.
El pipeline de recuperación combina pgvector (HNSW) para búsqueda semántica, pg_trgm (GIN) para expansión léxica e ILIKE como fallback — todo fusionado en un solo score ponderado. Extensiones de SQLite como sqlite-vec ofrecen búsqueda vectorial, pero carecen de:
Con más de 1,100 registros en dos tablas de memoria y creciendo, Hipocampo necesita una base de datos que escale sin sacrificar calidad de recuperación. PostgreSQL + pgvector no es "pesado" por capricho — es el stack mínimo viable para la precisión híbrida que BIRE y SSC exigen.
Hipocampo permite que agentes de IA aprendan de sus errores entre sesiones con un ciclo simple:
Ejemplo real: Un agente intenta flatpak install npm y falla. Guarda el error en Hipocampo: "npm es un gestor de paquetes de Node.js, no un paquete Flatpak. Usar npm directamente." La próxima vez que se intente el mismo comando, el agente encuentra este registro y aplica la solución de inmediato.
Con el tiempo, la base de conocimiento de errores crece orgánicamente. Cada fallo hace más inteligentes las sesiones futuras. Esto convierte a Hipocampo de un simple archivo en un sistema de aprendizaje continuo para agentes de IA.
Más allá del aprendizaje reactivo, Hipocampo v4.1 introduce reglas automáticas con disparadores que se activan antes de que el agente escriba una sola línea de código:
Cómo implementarlo:
Esto replica el hipocampo biológico: una pista parcial dispara la recuperación completa de la memoria — el cerebro no espera a que ocurra el error para recordar que duele.
A veces el agente rompe código que funcionaba bien — no repite un error viejo, crea uno nuevo. Hipocampo v4.2 implementa un ciclo inmune de 3 pasos que replica cómo el cuerpo genera anticuerpos:
Catálogo de archivos frágiles (precargado): header.php, conexion.php, utils.php, auth.php, db_connection.php — estos archivos tienen dependencias en cascada. Una sola edición incorrecta rompe decenas de páginas. Hipocampo incluye reglas de fragilidad para que los agentes sepan qué manejar con cuidado.
Antes de cada edición: search_hipocampo("trigger:regresion trigger:<archivo> trigger:<proyecto>") — aprende lo que otros agentes rompieron en este archivo antes de tocarlo.
Para activar este comportamiento, hay que instruir al agente. Se hace agregando reglas en su archivo de configuración:
| Agente | Archivo de configuración |
|---|---|
| OpenCode | AGENTS.md (raíz del proyecto) o ~/.opencode/AGENTS.md |
| Claude Code | CLAUDE.md o ~/.claude/CLAUDE.md |
| Cursor | .cursorrules |
| Windsurf | .windsurfrules |
| Cline | CLINE.md |
Ejemplo mínimo para AGENTS.md / CLAUDE.md:
💡 Tip: Para agentes nativos MCP (OpenCode, Claude Code), las tools de Hipocampo están disponibles directamente. Para otros, usa el endpoint HTTP o los scripts CLI.
📄 ¿Quieres que tu agente use TODO Hipocampo? Agrega las secciones de AGENTS_TEMPLATE.md a tu AGENTS.md / CLAUDE.md / .cursorrules — seguro de compartir, sin credenciales.
Los resultados de búsqueda BIRE pueden ser extensos (especialmente con el límite aumentado a 8000 caracteres). OpenCode trunca la salida de herramientas por defecto a 2000 líneas / 51 KB. Verás ...N bytes truncados... al final cuando ocurra. La salida completa se guarda en un archivo para lectura posterior, pero para evitar el truncamiento, agrega a tu opencode.json o ~/.config/opencode/opencode.jsonc:
Reinicia OpenCode para que el cambio surta efecto.
Soporte de SO: Ubuntu 22.04+, Debian 12+, Fedora 39+, Arch Linux, macOS (Homebrew) o Windows vía WSL2.
Para usar la búsqueda directamente desde la terminal:
Para inicializar el servidor MCP:
Operaciones de Memoria:
search_hipocampo(consulta, session_id?): Búsqueda semántica + léxica híbrida (auto-registra métricas). Filtro opcional por sesión.quick_hipocampo_search(consulta): Alias rápido para búsquedas.preload_context(ruta_proyecto, k=8): Extrae keywords del proyecto, busca memorias relevantes y devuelve resumen comprimido. Ideal para inicio de sesión.compress_hipocampo(consulta, k=5, method="hybrid", budget_ratio=1.0, include_metadata=False): Búsqueda + compresión híbrida con presupuesto de contexto. Auto-estima tokens y ajusta k dinámicamente. Tres métodos: "hybrid" (recomendado), "extractive" (más rápido, sin costo API), "llm" (máxima calidad).save_hipocampo(contenido, tipo, codigo, categorias, session_id?, force?, auto_link=False, nivel="episodica"): Guarda datos técnicos en memoria_vectorial. Soporta auto-dedup, auto-enlace y niveles jerárquicos.profile_hipocampo(resumen, extra, categorias): Guarda datos de perfil en memory_items.Grafo de Memoria (v4.0):
link_hipocampo(origen, destino, tipo_relacion, peso): Crea enlace dirigido entre recuerdos. Tipos: related, follow_up, part_of, references, similar, chain.unlink_hipocampo(id / origen+destino+tipo): Elimina enlaces del grafo.graph_hipocampo(nodo_id, profundidad=2): Árbol BFS desde un nodo raíz. nodo_id=0 muestra vista general.path_hipocampo(origen, destino, max_depth=5): Camino más corto BFS entre dos recuerdos.RAG de Código (v4.0):
index_project(ruta_proyecto, force=False): Indexa archivos de código como embeddings semánticos. Incremental — solo re-indexa archivos modificados. Soporta PHP, JS, TS, Python, SQL, HTML, CSS, JSON, YAML.search_code(consulta, k=5, lenguaje=""): Búsqueda vectorial en código indexado. Devuelve código real con ruta, lenguaje y líneas.Operaciones CRUD:
update_hipocampo(id, contenido?, tipo?, codigo?, categorias?): Actualiza un recuerdo existente. Regenera embedding si cambia el contenido.delete_hipocampo(id): Elimina un recuerdo permanentemente por ID.set_nivel_hipocampo(id, nivel): Promueve/degrada un recuerdo entre niveles jerárquicos (episodica, semantica, automatica).consolidate_hipocampo(dias_min=7, seco=True): Migra recuerdos episódicos antiguos a nivel semántico con compresión opcional.Autodiagnóstico y Reparación:
hipocampo_health(): Health check completo (PostgreSQL, API de embeddings, disco, extensiones, índice HNSW).hipocampo_auto_repair(): Repara problemas automáticamente (crea tablas, índice HNSW, reinicia PostgreSQL).Optimización de Rendimiento (Fase 2):
hipocampo_stats(): Métricas de rendimiento, latencia, y recomendaciones de optimización.hipocampo_tune(): Ajusta thresholds BIRE/SSC y pesos híbridos según uso real.Mantenimiento de Memoria (Fase 3):
hipocampo_dedup(fusionar): Detecta y fusiona memorias duplicadas (exactas + semánticas).hipocampo_checkpoint(seco): Checkpointing logarítmico para comprimir memorias antiguas.hipocampo_maintenance(): Ciclo completo de mantenimiento (reparar → dedup → checkpoint → tune).Decaimiento Temporal:
Olvido Activo y Tiering (v5.0):
decay_hipocampo(seco=True): Archiva memorias episodica antiguas a memoria_historica (almacenamiento frío). Protegidos: automatica, semantica, critico. Modo seco muestra qué se archivaría.hipocampo_budget(seco=True): Muestra distribución de memorias en tiers HOT/WARM/COLD. Cap del tier hot: 5000.restaurar_historica(id): Restaura una memoria fría del archivo al tier activo.contradicciones_hipocampo(id=None): Auditoría de contradicciones bajo demanda. Con ID: verifica una memoria. Sin ID: escanea todas.Watcher de Archivos (v5.0):
list_watch_dirs(): Lista los directorios monitoreados para auto-reindexación.add_watch_dir(path, patterns=["*.php","*.py","*.js"]): Agrega un directorio al watch list.remove_watch_dir(path): Elimina un directorio del watch list.reindex_now(path=None): Dispara reindexación inmediata de archivos modificados (o todos si no se da path).Webhooks (Watch):
watch_hipocampo(patron, webhook_url): Registra un webhook que se dispara en eventos save/update/delete cuando el contenido coincide con un patrón.unwatch_hipocampo(id): Elimina un webhook registrado.list_watches(): Lista todos los webhooks activos.La conexión a BD, configuración y generación de embeddings están centralizadas en el paquete hipocampo:
Todos los scripts en scripts/ importan de hipocampo.db en lugar de duplicar el boilerplate. El servidor MCP importa las funciones de búsqueda/salud/estadísticas/dedup/checkpoint directamente — sin llamadas subprocess.
Antes: Cada búsqueda MCP ejecutaba subprocess.run() → fork del intérprete Python → re-importar todo → conectar DB → generar embedding → ejecutar query → parsear stdout. ~200–500ms solo de overhead de proceso y serialización.
Ahora: Llamada directa a función en el mismo proceso. La DB connection pool, OpenAI client y módulos ya están cacheados. El overhead se reduce a microsegundos.
Para búsquedas individuales la diferencia es marginal (~200ms), pero para hipocampo_maintenance() antes ejecutaba 4 forks subprocess en serie — ahora es una llamada directa por fase, ahorrando ~1–2 segundos.
El servidor MCP ahora ejecuta las 16 herramientas como corutinas async en modo HTTP, y usa un pool de conexiones PostgreSQL en lugar de crear una conexión nueva por cada llamada:
Antes:
connect() en cada llamadasearch lenta congelaba el servidor para todos los clientes concurrentestoo many connections en la BDAhora:
init_pool(minconn=1, maxconn=10) crea un ThreadedConnectionPool al arrancar — las conexiones se reúsan, el handshake ocurre una sola vezasync def — I/O bloqueante (queries BD, API de embeddings) corre en asyncio.to_thread(), liberando el event loop para otras requests_PooledConnection devuelve las conexiones al pool automáticamente al llamar .close() — sin cambios en el callerImpacto: Requests concurrentes ya no se bloquean entre sí; el overhead de conexión PostgreSQL baja de ~10–50ms por llamada a casi cero.
Tests de Integración:
@pytest.mark.integration) arrancan el servidor en modo stdio y verifican tools/list, resources/list y una búsqueda realAntes:
NVIDIA_API_KEY o DB_HOST faltantes → el server arrancaba sin errores y fallaba con un críptico fe_sendauth / 401 recién en el primer queryexcept Exception: logger.error("msg: %s", e) — sin traceback, imposible saber si era error de BD, red o validaciónAhora:
validate_config() se ejecuta al arranque y logea warnings claros para cada variable faltante. init_pool() y get_conn() rechazan temprano con mensajes como "PostgreSQL connection incomplete: DB_HOST, DB_USER no configurados en .env"embedding_limiter (30/min — protege el costo de la API de embeddings), tool_limiter (60/min — protege PostgreSQL), watch_limiter (20/min). Los clientes reciben "⏳ Demasiadas solicitudes. Límite: 30 por 60s. Espera 12s."_tool_err() diferencia por tipo de excepción: psycopg2.Error → logger.exception() con traceback completo, ValueError / TypeError → logger.warning() (error del cliente), otros → logger.exception(). _fire_webhooks captura urllib.error.URLError por separadoImpacto: Los errores se detectan antes de llegar a la BD, los costos están limitados, y los logs son accionables — sabés al instante si es una mala configuración, un problema de red o un bug de código.
Consulte los manuales en la carpeta docs/ para información arquitectónica y configuraciones avanzadas.
Antes:
get_embedding() fallaba al primer timeout o rate limit de la API de embeddings — sin reintentossys.argv manual — sin --help, sin validación de tipos, interfaces inconsistentesAhora:
get_embedding() usa tenacity con wait_exponential(mult=1, min=1, max=30), 5 intentos, reintenta solo en RateLimitError/APITimeoutError/APIConnectionError/InternalServerError. No reintenta en AuthenticationError o BadRequestError. Cada reintento se loguea en nivel warningargparse con --help, argumentos tipados y nombres consistentes: hipocampo_mcp_server.py --http 8001.pre-commit-config.yaml con ruff lint+format (pre-commit) y pytest (pre-push). pyproject.toml configura ruff con line-length 120Impacto: El server tolera fallos transitorios de la API sin que el cliente vea errores. El CLI es autodocumentado. Cada commit se verifica antes de llegar a GitHub — no más tests rotos en main.
compress_hipocampo y Estructura basada en Symlinks (v3.9)Antes:
scripts/ eran copias independientes del repo — cada git pull requería sincronización manual, y archivos nuevos como hipocampo_compress.py no se propagabanscripts/ — sin separación de archivos del repohipocampo/ era también copia: load_config() buscaba .env en project_root/.env en vez de ~/.hipocampo/.env, cargando credenciales incorrectas (alex/hipocampo123)query_stats existía en esquema.sql pero nunca se creaba automáticamente — hipocampo_health reportaba DEGRADEDregister_vector(conn) sobre el _PooledConnection devuelto por get_conn() — psycopg2 lo rechazaba con TypeError, rompiendo todas las operaciones vectorialesinitialize, fallaban con Invalid request parametersAhora:
~/.hipocampo/ es el directorio canónico: repo/ (clon git), ~/.hipocampo/scripts/ → repo/scripts/ (symlink), ~/.hipocampo/hipocampo/ → repo/hipocampo/ (symlink). Scripts locales del usuario movidos a ~/.hipocampo/local_scripts/. git pull en repo/ actualiza todo automáticamente_find_env() en db.py carga .env en orden determinista: ENV_PATH → ~/.hipocampo/.env (config explícita del usuario) → project_root/.env (Docker/Fly). Sin más credenciales incorrectasensure_stats_table() se ejecuta al importar el módulo del MCP server — la tabla query_stats se crea automáticamente al iniciarhipocampo_compress.py con manejo explícito de excepciones: RateLimitError, APITimeoutError, APIConnectionError, APIStatusError → fallback inmediato a compresión extractiva. Otros errores → fallback con log de advertencia diferenciadoregister_vector(conn) eliminado de hipocampo_search.py, hipocampo_checkpoint.py, hipocampo_calibrate.py, mm_brain_tool.py — get_conn() ya registra el adaptador vectorial sobre la conexión realmcp[client] (stdio_client + ClientSession + initialize()) — handshake MCP 2025-03-26 correctoImpacto: Mantenimiento cero tras git pull. Errores transitorios de la API de embeddings degradan gracefulmente. Carga de configuración determinista y segura. Operaciones vectoriales confiables. Tests siguen el protocolo MCP oficial.
Consolidación, decaimiento y poda desatendidos — el sistema de memoria ahora se limpia solo. Tres mecanismos complementarios:
El servidor MCP puede ejecutar un ciclo de mantenimiento cada 24h (configurable). Desactivado por defecto — se activa con una variable de entorno:
| Variable | Default | Función |
|---|---|---|
HIPOCAMPO_AUTO_MAINTENANCE | false | Activa el scheduler en el lifespan HTTP |
HIPOCAMPO_AUTO_MAINT_INTERVAL_S | 86400 | Segundos entre ciclos de mantenimiento |
HIPOCAMPO_MAINT_MIN_AGE_DAYS | 7 | Edad mínima para consolidar episódica → semántica |
HIPOCAMPO_MAINT_DECAY_MIN_AGE_DAYS | 60 | Edad mínima para archivar episódicas sin acceso (olvido activo) |
Cada ciclo ejecuta: consolidación (episódica → semántica), decay (half-life 90d en enlaces + olvido activo de episódicas sin acceso), dedup merge y purga de access logs (>30d). Todas las protecciones intactas: automatica, semantica, critico y memorias enlazadas nunca se archivan.
Cada 50 saves (configurable con HIPOCAMPO_SAVE_TRIGGER_EVERY, 0 desactiva), un hilo en background ejecuta un ciclo liviano: consolidación + decay + purga — sin dedup merge (irreversible). El sistema se limpia solo en proporción a cuánto se usa, sin servicios externos.
Persistent=true (recomendado para PCs de escritorio)Para máquinas que se apagan de noche, cron pierde las ejecuciones programadas. Un timer systemd de usuario con Persistent=true recupera la tarea perdida apenas enciende la PC:
Instalar el timer semanal (domingo 03:00, catch-up al encender):
El CLI reutiliza _run_maintenance_cycle() del servidor MCP — exactamente el mismo código que el scheduler interno, cero duplicación de lógica.
decay_hipocampo nunca había corrido: la query de nivel memoria mezclaba timestamptz y text en un COALESCE (COALESCE(max(accessed_at), metadatos->>'date')) → error de PostgreSQL types timestamp with time zone and text cannot be matched. Corregido casteando (NULLIF(metadatos->>'date',''))::timestamptz. Mismo bug corregido en hipocampo_budget.min_age_days era ignorado: decay_hipocampo(dry_run=False, min_age_days=30) aceptaba el parámetro pero la Parte 2 (olvido activo) no tenía filtro de edad en el SQL — habría archivado episódicas de cualquier edad. Ahora el umbral de edad está parametrizado en la query.Optimizaciones aplicadas en Julio 2026 para reducir latencia y estabilizar thresholds:
get_embedding() ahora usa un caché LRU (128 entradas) — consultas repetidas con el mismo texto saltan la llamada a la API de embeddings, ahorrando ~450ms cada una. El cliente de OpenAI también se reutiliza entre llamadas.
SSC_TOP_K reducido de 20 → 15: menos resultados vectoriales por tabla = búsqueda más rápidahnsw.ef_search = 20: menor amplitud de búsqueda HNSW para vecinos aproximados más rápidos (default era 40)CONFIANZA_ALTA bajada de 70 → 60: salida temprana del pipeline SSC cuando los resultados vectoriales ya son buenosregister_vector cacheado por conexión para evitar introspección SQL redundantealpha reseteado de 0.6 → 0.5 (balanceado 50% vectorial + 50% léxico), vectorial_confidence_min de 0.75 → 0.70hipocampo_tune()) ahora limitado: alpha se mantiene entre 0.4–0.6, confidence entre 0.5–0.75register_vector() y detecta la incompatibilidad pgvector/PG17 (indam) con un mensaje claro de actualizaciónHipocampo incluye 103 tests unitarios cubriendo toda la lógica central e integración MCP:
| Archivo | Qué cubre |
|---|---|
tests/test_search.py | Expansión de consulta (stem map + sinónimos), fusión de scores con alpha dinámico, decaimiento temporal (5%/semana), formateo de resultados |
tests/test_autotag.py | Las 17 reglas de tags, 16 reglas de categoría, detección automática de memory_type |
tests/test_dedup.py | Similitud coseno (vectores de 1024 dim), lógica de detección de duplicados exactos y semánticos |
tests/test_checkpoint.py | Clasificación por escalas de edad, agrupación por proyecto, generación de resúmenes |
tests/test_mcp_integration.py | 6 tests de schema (registro de tools, anotaciones, parámetros, firma async) + 3 tests de integración en vivo (servidor stdio, mcp[client] SDK) |
tests/test_rate_limit.py | Rate limiter sliding-window: acquire/release, prune, stats, limiters por defecto |
tests/test_db.py | Validación de config: DB_HOST faltante, config de API de embeddings, cobertura completa |
Los tests se ejecutan automáticamente en cada push vía GitHub Actions en Python 3.11–3.13.
Si este proyecto te es útil, considera apoyarlo: