Privacy layer for enterprise APIs; refuses to start without a reviewed adapter jar per source
Copy the AI prompt to install this server into Claude Code, Cursor, or another agent β or use 1-click editor setup below.
One-click editor setup isnβt available for this listing yet β we donβt have a confirmed install command, and weβd rather show nothing than point your editor at the wrong package or host. Follow the projectβs own setup instructions, linked above.
Fail-closed privacy layer that pseudonymises enterprise API data for LLM agents and MCP clients.
Data Prism is an open-source privacy layer for Java/Spring teams putting LLM agents or MCP clients in front of internal APIs holding customer data. It pseudonymises personal data per privacy scope, refuses anything unclassified, and can keep a hash-chained audit trail.
Who it's for. Java/Spring platform and backend teams putting LLM agents or MCP clients in front of internal APIs that hold customer data. If nothing you run exposes personal data to a model, you don't need this.
Status: the walking skeleton and every slice through S9a are built, with 18
Maven submodules (19 Maven projects in the reactor counting the root
pom-packaged aggregator itself) and a passing test suite. The privacy
engine, correlation and consistency findings, parallel mTLS connectors,
embedded Hazelcast identity cache and read budget, an OAuth2 resource server
with session-derived PrivacyContext, audit and metrics are all real and
exercised end to end. The standalone server is the primary deployment
surface; the Spring Boot starter is the embedded option. A one-command local
Compose quickstart also exists: see "Try it" below. Two MCP tools ship
today, get_entity_context and compare_entity_sources β the other two
named in the design review, search_entity_data and
describe_entity_model, are not yet built (docs/tools.md "Not yet
built"). A durable, append-only, hash-chained audit sink and an offline
AuditChainVerifier ship as of 0.3.0, opt-in via
dataprism.audit.sink: hash-chained; the verifier catches an edit or
deletion inside a writer's chain, but cannot detect truncation of a writer's
most recent records or the deletion of a whole process boot's records, and
the trail does not resist an operator, or anyone else, who already has
write access to the file (docs/audit.md "What this does and does not
prove"). Not built: the re-identification operator surface (deferred past
V1 by decision, see docs/architecture.md#decisions-worth-knowing) and the
Elasticsearch connector and its search tools. See docs/plan/PLAN.md for
what is open.
An organisation wants an LLM to investigate live business data spread across several systems. Giving the model direct API access is not acceptable: those APIs carry personal and confidential data, each system represents the same entity differently, and raw identifiers let anything downstream correlate across sessions.
The obvious fix β redact everything sensitive β destroys the investigation. Once
three systems' names for one person are all [REDACTED], the model cannot tell
whether it is looking at one person or three.
It sits between the two and does two things that are easy to confuse:
It makes identity consistent. One subject gets one synthetic identity across
every source, derived deterministically from (scope, subject, namespace, algorithm version, key) β never random, never stored in plaintext, and
reproducible without the cache. The same person in three systems reads as one
person to the model.
It leaves the data inconsistent, and says so. If those three systems disagree about a name, the answer carries a finding that says they disagree. The platform never makes enterprise data look cleaner than it is. That distinction is the point of the project:
Identity representation becomes consistent. Underlying data inconsistencies become more visible, not less.
Pseudonyms are scoped. The same person in two different investigations gets two different synthetic identities, so nothing correlates across cases by accident.
Not an API gateway, not an ETL platform, not a master-data system, not an
identity provider, and not an entity-resolution engine β correlation requires a
key the sources already share, behind a documented SPI. It carries no business
domain: no Customer, Taxpayer or Employee type exists outside the example
application.
It is not anonymisation. Under GDPR Art. 4(5), pseudonymised data is still personal data. Sending Data Prism output to a third-party model is still processing, and still needs a lawful basis, a DPIA, and a transfer mechanism where the provider is outside the EU. The platform reduces exposure; it does not remove the obligation.
The fastest way to see a real MCP call answered by the real privacy engine β no local JDK, no Maven install, one command:
pulls the published ghcr.io/aindriub/data-prism-quickstart-<name> images
(pin one with QUICKSTART_IMAGE_TAG=0.3.1; run
docker compose -f compose.yaml -f compose.build.yaml up --build instead to
build every image from source) and brings up the standalone server, a
synthetic fixture API and a local HTTPS JWT issuer, proving an
agent-compatible get_entity_context call returns a pseudonymised response.
Walk through it in
docs/quickstart.md; connect your own agent client to
either that stack or a real deployment via
docs/agents/.
Once you have seen the demo, protect your own API: docs/quickstart.md ends
with a "What next" section pointing at
docs/protect-your-own-api.md, a YAML-only
walkthrough from a real JSON REST API to a working get_entity_context call.
The ghcr.io/aindriub/data-prism-server image listed there is published as a
multi-architecture manifest list covering linux/amd64 and linux/arm64,
each built and verified natively β docker run on Apple Silicon or any other
arm64 host pulls the arm64 image directly, no emulation required.
It is not a one-command install, on either architecture. docker run alone
yields a server that refuses to start: DataPrismContractValidator demands a
reviewed DataSourceAdapter bean for every configured source, and
DataPrismProperties.validate() demands a full deployment configuration (JWT
issuer/audience/JWKS, caller-claim mappings, security policy, HMAC key
reference, audit sink, metrics sink, Hazelcast topology). Neither ships in the
image. Two things an operator must supply themselves before it serves
anything:
DataSourceAdapter (and IdentityResolver) jar for each
API you are protecting, mounted onto the image's loader path.dataprism.* vocabulary.docs/configuration.md is the authoritative,
complete contract for both. The "Try it" section above is a local Compose
fixture for evaluation, not this image or that configuration.
The full set of user docs is also published, rendered and searchable, at https://aindriub.github.io/data-prism/.
Each row links the site page and the repo file it is built from.
No reviews yet β be the first to share how this listing worked for you.
Showcase your server listing on GitHub or your project documentation. Embed this dynamic SVG badge to highlight official listing status and live engagement.
[](https://allmcps.com/mcp/data-prism)<a href="https://allmcps.com/mcp/data-prism"><img src="https://allmcps.com/api/badge/data-prism?style=directory" alt="Data Prism on AllMCPs" /></a>