The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Facesign MCP listing page.
Import and initialize a client using an integration token
Make a request to any Facesign API endpoint.
A flow is a directed graph of nodes that defines a session. Sessions can serve different purposes — identity verification, data collection, authorization, analysis, or any combination. Every flow must start with a START node and end with one or more END nodes. Nodes connect to each other via outcomes — each outcome points to the id of the next node. All defined outcomes must be connected to other nodes; no outcome can be left unlinked.
The entry point of every flow. Each flow must have exactly one START node. It has a single outcome that points to the first node in the flow.
The terminal node of a flow. A flow can have multiple END nodes (e.g., one for success path, one for failure path). It has no outcomes.
The avatar speaks to the user using the prompt text and routes the flow based on the user's response. The session stays on this node until one of the outcome conditions matches the user's response, so a single conversation node can facilitate a multi-turn dialog with the user until the desired information is gathered.
The prompt field supports two modes:
Direct speech — Tell the avatar exactly what to say:
"Say: Hi, how are you doing?"
Goal-driven behavior — Describe what the avatar should achieve during the conversation:
"Chat with the user and find out how well they understand medicine."
When using goal-driven prompts, always include realistic exit conditions in the outcomes, since the dialog could otherwise continue indefinitely. Examples of exit conditions:
Uses conditional outcomes — each outcome has a condition (a natural language description of what the user's response should match) and a targetNodeId pointing to the next node.
Set doesNotRequireReply: true for nodes at the end of the flow where the avatar delivers a final message and no user response is needed (e.g., "Thank you for the conversation, goodbye!"). Typically used before an END node.
Outcomes use FSConditionalOutcome:
id — unique identifier for the outcometargetNodeId — the node to navigate tocondition — natural language description of the matching criteriaUnconditional transition: When the avatar should say a phrase and move to the next node regardless of the user's response, use a single outcome with an empty condition (""). This creates an unconditional transition — the avatar delivers the message, and whatever the user replies (or even if they don't), the flow proceeds to targetNodeId.
Direct speech example:
Goal-driven example:
End-of-flow example:
Requests camera and/or microphone permissions from the user before proceeding. Use this node when you need to display a custom prompt explaining why permissions are needed, or when you want to handle the denied case with a specific flow path.
If the flow does not contain any PERMISSIONS node, permissions will be requested automatically.
permissions.camera — request camera accesspermissions.microphone — request microphone accessprompt — optional message the avatar will say (uses direct speech mode, e.g., "Say: Could you please enable your microphone so I can hear you."). If the site already has permanent permissions granted, the avatar will not say this phrase and the flow will continue directly via the permissionsGranted outcome.Outcomes:
permissionsGranted — user granted the requested permissions (or they were already granted)permissionsDenied — user denied the permissionsChecks whether the user in front of the camera is a real person or a deepfake. The node analyzes the video feed to perform liveness detection.
Prerequisites: The user must have camera access granted before this node. Additionally, liveness detection requires several seconds of video recording for analysis. Do not place this node immediately after a PERMISSIONS node — instead, add a CONVERSATION node in between to give the system time to accumulate video data for analysis.
Recommended flow order:
PERMISSIONS → CONVERSATION → LIVENESS_DETECTION
Outcomes:
livenessDetected — the user is a real persondeepfakeDetected — a deepfake or spoofing attempt was detectednoFace — no face was detected in the camera feedDisplays a UI for the user to enter their email address without any verification or confirmation step. Use this node when you simply need to collect an email from the user as data input. Note: if your goal is to collect AND verify an email via OTP, use TWO_FACTOR_EMAIL instead — it handles email collection internally and does not require a preceding ENTER_EMAIL node.
prompt — optional phrase the avatar will say at the moment the input field appears (e.g., "Could you please enter your email?").Outcomes:
emailEntered — user submitted their emailcanceled — user canceled the email entryValidates data collected during the session and routes the flow based on the result. Uses a validation object to specify which field to check, what action to perform, and an optional expected value.
validation.field — the data field to validatevalidation.action — the validation action to performvalidation.value — optional expected value for comparisonUses conditional outcomes (same as CONVERSATION node) to branch the flow based on the validation result.
Opens a document scanning UI powered by Microblink. The user can scan identity documents using their camera. Extracted data (name, date of birth, document number, etc.) becomes available in the session report.
scanningMode — determines how to scan the document:
"single" — scan only one side of the document"automatic" — automatically determine how many sides need to be scannedallowedDocumentTypes — array of document types the user can scan (e.g., "passport", "id", "dl", "residence-permit", "visa", etc.)showTorchButton — show flashlight toggle (default: true)showCameraSwitch — show front/back camera toggle (default: true)showMirrorCameraButton — show mirror camera button (default: true)Outcomes:
scanSuccess — document was scanned successfullyuserCancelled — user canceled the scanscanTimeout — scan timed outPerforms biometric face recognition to identify the user. Compares the user's face against previously registered faces to determine if they are a known or new user.
Prerequisites: Requires camera access. Like LIVENESS_DETECTION, needs several seconds of video for analysis — do not place immediately after a PERMISSIONS node. Add a CONVERSATION node in between to accumulate video data.
Outcomes:
recognized — the user was matched to a known facenewUser — the user's face was not found in the database (new user)noFace — no face was detected in the camera feedPerforms 1:1 biometric face matching. Captures the user's face and compares it against a reference image to verify their identity.
Prerequisites: Requires camera access. Like LIVENESS_DETECTION, needs several seconds of video — do not place immediately after a PERMISSIONS node.
captureInstructions — optional text instructions shown to the user during capturerequireLivenessChallenge — require an active liveness challenge during capturerequireAILivenessCheck — require AI-based liveness verificationreferenceImageKey — key identifying the reference image to compare againstsimilarityThreshold — minimum similarity score (0–1) required for a matchenableSound — enable audio feedback (default: true)enableHaptics — enable haptic feedback (default: true)Outcomes:
passed — face matched the reference imagenotPassed — face did not matchcancelled — user canceled the scanerror — an error occurred during scanningSends a one-time password (OTP) to the user's email address and verifies the code they enter. This node manages its own sub-flow for collecting the email: if no email was provided via providedData or found in publicRecognition, the node will automatically prompt the user to enter their email. There is no need to add an ENTER_EMAIL node before this one — email collection is handled internally.
otpLength — number of digits in the OTP (4–8, default: 6)expirySeconds — how long the OTP is valid (default: 300 / 5 minutes)maxAttempts — maximum verification attempts (default: 3)resendAfterSeconds — minimum delay before "Resend" button is enabledshowUI — show on-screen toast notification (default: true)emailTemplate — optional custom email templateOutcomes:
verified — user entered the correct OTPdelivery_failed — OTP could not be deliveredfailed_unverified — user exhausted all attempts without verifyingcancelled — user canceled the verificationerror — an error occurredCompares faces from two different sources to verify they belong to the same person. Unlike FACE_SCAN (which captures a photo and compares it against a single reference), this node takes already-existing images from different session sources and compares them.
Available sources:
sessionVideo — a frame captured from the live video feed during the sessionfaceScan — a higher-quality photo from a FACE_SCAN node (oval capture). Uses the most recent completed FACE_SCAN node in the session.providedData — an image URL from a providedData field. Requires providedDataKey to specify which field contains the URL.documentPhoto — a photo extracted from a scanned document (DOCUMENT_SCAN node). Uses the most recent completed DOCUMENT_SCAN node in the session.Configuration:
sourceA — first image sourcesourceB — second image sourcesimilarityThreshold — optional minimum similarity score (0–1) required for a matchOutcomes:
match — faces from both sources matchnoMatch — faces do not matchimageUnavailable — at least one image could not be obtained (e.g., no person detected on camera, missing providedData field, no photo on document, etc.)Example — compare live video with a photo from providedData:
Example — compare face scan capture with document photo:
Same as TWO_FACTOR_EMAIL but sends the OTP via SMS. This node manages its own sub-flow for collecting the phone number: if no phone number was provided via providedData or found in publicRecognition, the node will automatically prompt the user to enter their phone number. There is no need to add a separate phone collection node before this one.
smsTemplate — optional custom SMS template (instead of emailTemplate)You can ask the backend to extract structured data from the session transcript using an LLM. Configure it via SessionSettings.extractionSchema when creating a session; results are returned in SessionReport.extractedData.
Each entry in extractionSchema describes one field:
fieldName — key under which the extracted value will appear in extractedData.type — "string", "number", "boolean", or "date". "date" is returned as an ISO 8601 string.description — natural-language hint for the LLM describing what to look for. This is the primary signal for extraction, so be specific.enum — optional list of allowed string values. When set, the LLM normalizes free-form answers (e.g., "yeah", "sure") into one of the listed values.Every field is treated as optional: if the transcript does not contain the data, the value in extractedData will be null. There is no required flag.
Resulting shape on the session report: