Find open models and get download links with the SHA-256 each file must have.
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.

Hologram Live is a local-first module host for the Hologram ecosystem. This repository produces two independent products:
hologram binary, containing the CLI and background service.hologram as a managed sidecar and embeds the shared .holo application executor.The current desktop experience provides a Console dashboard, multi-thread Chat with archiving, content-addressed Files, watched .holo Applications, and module discovery in a responsive dark/light interface. A Cmd/Ctrl+K command palette reaches every action, and text size is adjustable with Cmd/Ctrl +/-/0. Chat routes through a configurable inference engine: the default echo engine repeats your message, while weightc (one-shot CLI over imported .wcpu artifacts) and Ollama-compatible HTTP endpoints serve real model completions.
The preparation step builds the server sidecar before Tauri opens. The desktop app can:
.holo archives,
and inspect verified archive metadata;Build an installable desktop bundle with:
The desktop UI's colours, type, and marks derive from the Hologram brand kit at a pinned revision; apps/desktop/BRANDING.md explains how to update the brand everywhere in three steps.
Run the service in the foreground instead with:
Every server advertises http://127.0.0.1:11435 by default and generates a
256-bit membership secret in its state directory at cluster.token. The file
is reused across restarts and is owner-only on Unix. To form a multi-host
cluster, securely give every node the same secret (at least 32 bytes), advertise
the origin other nodes can reach, and seed a new node with any live member:
The joining node heartbeats immediately, learns the live membership set, and
then heartbeats those peers directly. Failed seeds remain eligible for retry;
members disappear from /api/v1/nodes after their TTL. Join messages use a
short-lived keyed proof over the exact payload, so the cluster secret and the
separate user authentication token are never sent over the wire. Non-loopback
origins must use HTTPS.
The default configuration and local endpoint are:
Open the endpoint for the built-in status page. API documentation is available at:
/docs is the self-hosted Scalar reference; /openapi.json is the generated OpenAPI document. Native clients use the versioned Protobuf/gRPC service on the same endpoint.
A path no route claims is answered 404 with the daemon's error envelope (LIVE_NOT_FOUND). A caller that sends content-type: application/grpc gets gRPC UNIMPLEMENTED instead, as any gRPC server answers an unknown service.
Global --json is supported by every CLI command and may appear before or after the subcommand. A successful command writes one JSON value to stdout, including lifecycle actions, downloads, generated files, accepted mutations, and run --output-format text. Diagnostics remain on stderr, while a runtime failure writes a JSON object with code and message to stdout and exits nonzero. This makes the complete CLI safe to compose with jq:
Help and shell-completion text retain Clap's human-readable format.
Create a thread, copy its returned ID, and send a message:
chat send records the user message and the assistant response as one persisted exchange. Threads retain separate histories and can be resumed from the desktop app.
The response comes from the inference engine selected in live.toml:
The default echo engine repeats the user message; it needs no model and no external process. The weightc engine shells out to weightc ask <artifact-dir> <prompt> --json against an imported .wcpu artifact. ollama proxies /api/generate, while vllm uses the OpenAI-compatible /v1/completions and /v1/models endpoints with native streaming and optional authentication from VLLM_API_KEY.
llamacpp loads a local GGUF model in-process and streams decoded pieces with exact token counts. Set model_path directly, or import a GGUF file and put its returned blake3:... id in default_model. It is off by default because it builds native C++ code and gives model execution the daemon's crash boundary. Build it with cargo build --release --features llamacpp; use llamacpp-metal or llamacpp-cuda for the corresponding GPU backend. These builds require CMake, Clang, and a C++ compiler.
candle is the Rust-native local GGUF alternative. The initial adapter supports Candle's quantized Llama-family implementation, requires a matching Hugging Face tokenizer.json, streams native deltas, and reports exact token counts. Set model_architecture = "llama"; use --features candle for CPU, candle-metal for Apple GPUs, or candle-cuda for NVIDIA GPUs. Candle is not a universal GGUF dispatcher: unsupported model families fail during startup instead of being guessed.
burn uses Tracel's Burn-LM Llama implementation on CPU. It requires a Burn named-MPK checkpoint, the matching Llama 3 tokenizer.model, and one of llama3.2-1b, llama3.2-3b, llama3.1-8b, or llama3-8b in model_architecture. Build with --features burn. Burn generation is currently buffered, so streaming API responses honestly report x-hologram-stream: emulated; GGUF files are not Burn checkpoints.
Run the live acceptance gate against real engines before releasing an inference build. It starts an isolated Hologram server and checks model discovery, buffered and native-streaming OpenAI/Ollama requests, usage, cancellation, and llama.cpp context overflow:
VLLM_API_KEY is forwarded when set. Set HOLOGRAM_BIN to test an existing binary, or let the script build the required feature set.
With resident_sessions = true, the weightc engine instead keeps a supervised weightc enter --jsonl process per conversation, so turns reuse the live KV context instead of replaying a transcript, and only the new message crosses the wire each turn. Sessions are LRU-capped by max_resident_sessions; a crashed session is reported as a typed error and lazily respawned (starting fresh context) on the next turn. This mode needs a weightc build with enter --jsonl support. Models are managed with:
Threads can be archived instead of deleted. Archived threads keep their messages and their updated_at_millis, but drop out of the default listing:
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/hologram-model-hub)<a href="https://allmcps.com/mcp/hologram-model-hub"><img src="https://allmcps.com/api/badge/hologram-model-hub?style=directory" alt="Hologram Model Hub on AllMCPs" /></a>