The SDK runs in your application process and requires a compatible native library for its operating system and CPU. Use the distribution supplied for your deployment; do not assume that a successful package install proves the native library loaded.
Python: package minimum 3.9; individual integrations can require newer Python. Use Python 3.10 or newer for the MCP walkthrough. The declared LangChain extra targets langchain-core 0.3.x.
TypeScript: run on Node.js with native-library support. Browser and edge-worker runtimes are not interchangeable with Node. Confirm the Node version and OS/CPU combination supported by your supplied release before deployment.
Install the framework and provider packages your application uses. They are not all installed by the base SDK.
The local quickstarts below need no model service or provider key. Use a fresh test directory and fictional input. They exercise real policy enforcement with a small model stand-in.
Getting started
Create a policy
Save this as policy.json. Both SDKs read the same policy format. It allows model calls, blocks a synthetic SSN pattern, and refuses tool names beginning with delete_.
With the Python package installed, you can instead generate a broader starter policy. The command writes the file itself; do not redirect its printed instructions into the JSON file.
Policy CLI
if-policy init -o starter-policy.json
if-policy validate policy.json
if-policy show policy.json
import inferencefort
from inferencefort import PolicyViolation
from langchain_core.language_models.fake_chat_models import FakeListChatModel
class CountingModel(FakeListChatModel):
calls: int = 0
def _call(self, *args, **kwargs):
self.calls += 1
return super()._call(*args, **kwargs)
model = CountingModel(responses=["Allowed response"])
print(model.invoke("Say hello").content)
try:
model.invoke("Synthetic SSN: 123-45-6789")
except PolicyViolation:
print("Blocked before model execution")
else:
raise RuntimeError("Expected a policy block")
assert model.calls == 1, "The blocked prompt reached the model"
status = inferencefort.governance_status()
assert status.get("initialised"), "Governance did not initialise"
print("Verified: one allowed call, one blocked call")
Run and verify
Run
python quickstart.py
Expected output includes Allowed response, Blocked before model execution, and the verification message. The counter must remain one. Replace the stand-in with your supported provider only after this succeeds; install that provider integration and configure its endpoint and credentials separately.
This executable JavaScript example uses the TypeScript SDK in ESM mode. Explicit wrapModel avoids import-order ambiguity. The model stand-in uses the same model-call surface intercepted in Vercel AI SDK integrations.
Install ai and @ai-sdk/openai, configure your provider key through your secret manager, and use this model with generateText. This makes a real provider request and may incur usage charges.
Python installs lazy interception through the packaged startup hook. Explicitly importing inferencefort or using if-run app.py also activates the hooks. Export configuration before starting the process. Python started without normal site initialization does not load the startup hook.
For Node applications that need automatic interception, load the registration hook before application imports:
Calling activate() after strict ESM imports is not a substitute for the preload hook. Alternatively, use explicit wrappers for the surfaces your application calls.
Surface
Python
TypeScript / Node
Verify
LangChain / LangGraph models
Lazy framework interception
Preload or explicit framework wrappers
Invoke and streaming paths used by your app
Vercel AI SDK
Not applicable
Preload or wrapModel
Model calls and tool execution separately
LiteLLM
Completion and async completion
Not applicable
Actual provider routing
Raw OpenAI / Anthropic
Supported client interception
Preload or wrapOpenAI / wrapAnthropic
The specific API method, including streaming, used by your app
CrewAI tools
Supported tool interception
Not applicable
Agent-loop tools as well as direct calls
MCP
Client tool calls and server trace continuation
Preload or MCP module wrappers
Client/server exchange below
Custom HTTP clients / custom tool loops
Do not assume interception. A supported model client does not prove a custom tool loop is governed.
Deliberate blocked action and coverage diagnostics
Other model providers can be reached through supported framework integrations. This does not imply interception of every provider’s raw SDK or every new API method. Recheck your exact paths when upgrading frameworks. Streaming input is evaluated before provider execution; output controls cannot undo effects or content already delivered.
Getting started
Verify enforcement
Run the local quickstart for your language.
Configure the actual framework, provider, and tool path your application uses.
Run one permitted call and one deliberately prohibited call.
Confirm the prohibited call never reaches its provider or tool handler.
Review status, findings, and audit events. A successful request or an empty findings list alone does not prove coverage.
Integrations
MCP tools
Complete the Python quickstart first to create the environment and export the policy configuration. For the Node client, also install the SDK from the TypeScript quickstart.
Keep the example policy active. This local server returns fixture text and never deletes anything. Both clients below must print a status result and a block; neither should receive Unexpected execution. Run each client from the directory containing the server file. Stdio transports may filter inherited environment variables, so these examples explicitly pass the absolute policy path to the server without forwarding unrelated credentials.
Python 3.10+ · MCP dependency
pip install "inferencefort[mcp]"
mcp_server.py
import inferencefort
from mcp.server.fastmcp import FastMCP
server = FastMCP("documentation-example")
@server.tool()
def lookup_status() -> str:
"""Return an example status."""
return "Example status: ready"
@server.tool()
def delete_record() -> str:
"""Fixture only; never deletes data."""
return "Unexpected execution"
if __name__ == "__main__":
server.run(transport="stdio")
mcp_client.py
import asyncio
import os
import sys
import inferencefort
from inferencefort import PolicyViolation
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
async def main():
params = StdioServerParameters(
command=sys.executable, args=["mcp_server.py"],
env={"INFERENCEFORT_POLICY_FILE": os.environ["INFERENCEFORT_POLICY_FILE"],
"INFERENCEFORT_FAIL_MODE": "fail_closed"},
)
async with stdio_client(params) as (reader, writer):
async with ClientSession(reader, writer) as client:
await client.initialize()
print(await client.call_tool("lookup_status", {}))
try:
await client.call_tool("delete_record", {})
except PolicyViolation:
print("Blocked delete_record before transport")
else:
raise RuntimeError("Expected a blocked tool call")
asyncio.run(main())
Run Python client
python mcp_client.py
TypeScript SDK client against the same server
Keep the Python environment active so python resolves to the interpreter with MCP installed. Wrappers accept module namespaces, not constructed clients, and return a boolean indicating wrapping success.
Node MCP dependency
npm install @modelcontextprotocol/sdk
mcp_client.mjs
import * as mcpClient from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';
import { wrapMcpClient, PolicyViolation } from '@inferencefort/ai-sdk';
if (!wrapMcpClient(mcpClient)) throw new Error('MCP client wrapping failed');
const client = new mcpClient.Client({ name: 'documentation-client', version: '1.0.0' });
try {
await client.connect(new StdioClientTransport({ command: 'python', args: ['mcp_server.py'],
env: { INFERENCEFORT_POLICY_FILE: process.env.INFERENCEFORT_POLICY_FILE,
INFERENCEFORT_FAIL_MODE: 'fail_closed' } }));
console.log(await client.callTool({ name: 'lookup_status', arguments: {} }));
try {
await client.callTool({ name: 'delete_record', arguments: {} });
throw new Error('Expected a blocked tool call');
} catch (error) {
if (!(error instanceof PolicyViolation)) throw error;
console.log('Blocked delete_record before transport');
}
} finally {
await client.close();
}
Wrap the low-level server module before constructing the server and registering its handlers. Continue your normal MCP server setup after this fragment. Inbound trace continuation does not replace authentication or server authorization.
Node server initialization · fragment
import * as mcpServer from '@modelcontextprotocol/sdk/server/index.js';
import { wrapMcpServer } from '@inferencefort/ai-sdk';
if (!wrapMcpServer(mcpServer)) throw new Error('MCP server wrapping failed');
const server = new mcpServer.Server(
{ name: 'example-server', version: '1.0.0' },
{ capabilities: { tools: {} } },
);
// Register your tool handlers and connect your transport here.
Configuration
Controls
Model access, content rules, configured-destination restrictions, and tool policies can be evaluated locally from policy. Configured detectors, session services, and shared budgets may require network requests. The controls available depend on your account, deployment, and integration coverage.
Control
Configuration and behavior
Model access
Allow or deny models by policy. Include an explicit allow policy for intended usage; unmatched access is denied when policy is loaded.
Content rules
Substring or regex rules with block or flag actions, scoped to input, output, or both. Output decisions occur after execution.
Egress
Provider-to-destination allow-lists check the configured destination. Matching is host-anchored; loopback is allowed by default. Pair with network controls.
MCP servers
An empty allow-list is unrestricted; stdio/in-process calls are local. Configure remote servers explicitly.
Tool policy
blocked_tools and approval_tools accept tool-name patterns. An approval-required verdict stops the call; build the human review and authorized retry into your application.
PHI and detectors
Enable the intended scanners and redaction separately. Install/configure the required redactor or detector endpoint and check capability gaps. Detection does not guarantee every sensitive value is found.
Session detection
Account-dependent controls use session context. Findings and configured actions determine blocking, flagging, or approval; do not assume every finding is removed from returned text.
Budgets
Shared daily caps need the control plane. Local-only operation does not provide a shared ledger.
Policy fragments
Merge the relevant fragment into your reviewed policy; these are not standalone complete policies.
Use enforcement_mode: "monitor" to evaluate without refusing policy violations, and "enforce" to block. Review monitor findings before enabling enforcement. Do not use monitor mode for the blocked-call walkthroughs.
For PHI in Python, install inferencefort[hipaa] and the selected spaCy model, then enable redaction in policy. Node deployments must configure their supported redaction service when required, including INFERENCEFORT_PHI_URL. Confirm capability status rather than assuming equivalent redaction coverage from package installation.
Configuration
Connect a control plane
Use the control-plane address supplied for your hosted workspace or your own deployment. The example address below is a placeholder, not a working service. Set keys through your deployment secret manager; never commit them into application code, policy files, or documentation.
Hosted or self-hosted configuration
export INFERENCEFORT_API_URL="https://control-plane.example.com"
# Set INFERENCEFORT_KEY through your secret manager before starting the app.
The default API address is http://localhost:8080; setting a key does not change it. A provider key is separate from your InferenceFort key. Verify connectivity, policy availability, and the workspace’s enforcement mode before relying on a blocked action.
Backend policy refreshes on a server-controlled TTL. Changes are not instantaneous. Existing cached rules continue to apply during refresh failures; verify the active policy before judging a change. Audit delivery is asynchronous and should not be treated as proof of durable receipt merely because a call returned.
Local policy plus backend policy
A configured local file can remain active alongside backend policy; either layer can tighten restrictions. For an agent fleet, set INFERENCEFORT_POLICY_DIR to a directory of JSON files declaring agent_id, and optionally system_id to disambiguate systems. Do not put backend-managed pricing, budget, detector, classification, or kill-switch settings into new per-agent files. Validate each file before deployment.
Identity & sessions
Identity & isolation
Set identity from your authenticated server context before the agent runs. Do not trust a tenant or user identifier supplied in an arbitrary request body. Session state is keyed by customer and conversation, so reuse the conversation ID across requests in the same conversation and separate customers explicitly.
Identity is optional unless policy requires it. Missing an explicit conversation ID causes a fallback ID to be pinned to the current context; later calls in that context reuse it. The fallback does not automatically join separate HTTP requests into one conversation.
Python contextvars propagate through await, into tasks created under the context, and through asyncio.to_thread. Raw executor submissions need explicit propagation. Set identity before spawning work, not after the call or in a sibling task.
Python raw executor · fragment
from inferencefort.core.context import bind_context
await loop.run_in_executor(pool, bind_context(lambda: agent.invoke(message)))
Keep request scope outside the graph run so child nodes inherit it. For LangGraph, pass invocation configuration explicitly as a keyword in Python. For Node worker threads or separate processes, transfer correlation context explicitly rather than expecting AsyncLocalStorage to cross that boundary.
Identity & sessions
Cross-agent traces
MCP integrations can propagate trace context through request metadata. For your own HTTP or queue transport, inject the context on the caller and extract it on the receiver. These fragments assume authenticated transport and an existing governed caller scope.
Python correlation · application fragment
from inferencefort import inject_context, extract_context, remote_context
headers = {}
inject_context(headers)
# Send headers with your authenticated request.
# Receiver: establish its own authorized customer identity first.
with remote_context(extract_context(incoming_headers)):
result = agent.invoke(message)
Node correlation · application fragment
import { injectContext, extractContext, withRemoteContext } from '@inferencefort/ai-sdk';
const headers = {};
injectContext(headers);
// Send headers with your authenticated request.
// Receiver: establish its own authorized customer identity first.
const result = await withRemoteContext(
extractContext(incomingHeaders), () => agent.invoke(message),
);
Correlation headers are not authentication. The receiver must establish its own customer identity. Cross-process session continuation depends on compatible trust configuration; trace linkage alone does not prove session state is shared. Coordinate shared session configuration across participating deployments and verify an end-to-end trace.
Configuration
Offline deployment
No-key local policy supports static model access, content rules, egress, MCP server restrictions, and blocked/approval tools. It does not enable account-dependent session detection, including the lethal-trifecta guard, value provenance, runtime tool classification, or session-risk classification. It also does not provide shared budgets or the hosted audit trail.
Configured customer detectors, redaction services, and SIEM forwarding may still make network requests. “No InferenceFort account” and “no outbound network” are different deployment properties. For air-gapped use, keep required services within the permitted boundary, review every configured endpoint, and verify the supplied deployment’s supported capabilities.
Restart after changing a static policy file. To propose policy from observed audit traffic, use if-policy learn audit.ndjson; review the proposal and validate it before switching from monitor to enforcement. Audit files can contain sensitive content and must follow your organization’s access and retention rules.
Operations
Troubleshooting
Inspect status after a deliberate governed call: initialization may be lazy. Check both initialization and capability gaps. Initialization confirms a loaded policy, not coverage of every framework or tool path.
import { governanceStatus } from '@inferencefort/ai-sdk';
const status = governanceStatus();
console.log('Initialised:', status.initialised);
console.log('Capability gaps:', status.capabilityGaps);
Symptom
Check
Installed, but no block
Configuration before startup, loaded policy, enforce mode, matching rule, and the actual intercepted call path.
Not initialized
Policy path or control-plane connectivity, native library availability, and release compatibility with this runtime.
ESM model calls pass through
Preload registration before imports or explicitly wrap the model.
Model governed, tool ungoverned
Confirm the tool integration separately; model interception does not wrap arbitrary application actions.
Rule misses expected text
Check input/output scope and scanned_scopes. System prompts and assistant history require the relevant scanning configuration.
Redaction or detector unavailable
Review capability gaps, required dependencies, endpoint reachability, and enabled configuration.
Wrong user or session
Set authenticated identity inside the request scope before creating child tasks; verify trace attribution.
No hosted audit events
Check account configuration, connectivity, content-logging settings, and asynchronous delivery. No-key mode has no hosted audit trail.
For focused troubleshooting, enable stage logging temporarily. Restrict access to diagnostic output, inspect it for sensitive content before sharing, and disable it when finished.
Choose one mode for your deployment. Cached local rules are still evaluated; a remote outage does not automatically disable all enforcement. Configured detectors and other services can have their own failure settings, which also need review.
Mode
Important behavior
fail_open (default)
Availability-oriented fallback when governance cannot obtain a required answer. Calls may proceed without the unavailable control.
fail_closed
Refuses calls when a required policy cannot be loaded or a budget check fails. Test the resulting application errors before enabling.
fail_cached
For budget outages, reuses the last successful budget answer when available. Without a cached budget answer, the call can proceed. It is not a universal replay of every previous verdict.
Catch PolicyViolation to handle a refused action. Rethrow unrelated exceptions; provider and application failures remain distinct. For approval-required actions, present the proposed action for authorized review and implement the supported approval/retry flow. Never treat merely catching an exception as approval.
Operations
Environment variables
Variable
Purpose
INFERENCEFORT_KEY
Workspace SDK credential; inject through a secret manager.
INFERENCEFORT_API_URL
Control-plane base URL. Default http://localhost:8080.
INFERENCEFORT_POLICY_FILE
Path to a reviewed static policy bundle.
INFERENCEFORT_POLICY_DIR
Directory of per-agent JSON policies.
INFERENCEFORT_FAIL_MODE
fail_open, fail_closed, or fail_cached.
INFERENCEFORT_CUSTOMER_ID
Deployment-level customer identity where appropriate; use request scope for multiple customers.
INFERENCEFORT_PHI_URL
Node delegated redaction endpoint when required by your deployment.
INFERENCEFORT_MCP_PROPAGATE
Set to 0 to disable automatic MCP correlation propagation.
INFERENCEFORT_SHARED_TAINT
Shared session capability configuration; coordinate across participating agents.
Enable and choose the destination for local stage diagnostics.
This is the configuration used in these guides, not a promise that every framework or release accepts every optional setting. Keep the SDK, native library, and framework versions aligned with your supplied release.