InferenceFortDocs

Introduction

Enforce policy on your AI agents’ model and tool calls. Start with a local policy, verify a blocked action, then connect your deployment.

Start with proof. The quickstarts use synthetic data and a model stand-in. No provider key or running model service is needed.

Choose your SDK

Your integration path

  1. Create a policy

    Define which model calls and tool actions are permitted.

  2. Verify enforcement

    Confirm a prohibited request never reaches the model or tool.

  3. Govern your tools

    Test an allowed and blocked MCP action.

  4. Connect your deployment

    Configure the control plane and review operational coverage.

Before production

Review integration coverage, request identity, and outage behavior for the exact paths your application uses.

Requirements

The SDK runs in your application process and requires a compatible native library for its operating system and CPU. Use the distribution supplied for your deployment; do not assume that a successful package install proves the native library loaded.

  • Python: package minimum 3.9; individual integrations can require newer Python. Use Python 3.10 or newer for the MCP walkthrough. The declared LangChain extra targets langchain-core 0.3.x.
  • TypeScript: run on Node.js with native-library support. Browser and edge-worker runtimes are not interchangeable with Node. Confirm the Node version and OS/CPU combination supported by your supplied release before deployment.
  • Install the framework and provider packages your application uses. They are not all installed by the base SDK.

The local quickstarts below need no model service or provider key. Use a fresh test directory and fictional input. They exercise real policy enforcement with a small model stand-in.

Create a policy

Save this as policy.json. Both SDKs read the same policy format. It allows model calls, blocks a synthetic SSN pattern, and refuses tool names beginning with delete_.

policy.json
{
  "version": "docs-example-1",
  "enforcement_mode": "enforce",
  "policies": [
    {
      "name": "allow-example-models",
      "effect": "allow",
      "models": [
        "*"
      ],
      "enabled": true
    }
  ],
  "rules": [
    {
      "name": "example-ssn",
      "kind": "regex",
      "pattern": "\\b\\d{3}-\\d{2}-\\d{4}\\b",
      "action": "block",
      "applies_to": "input",
      "enabled": true
    }
  ],
  "mcp_policy": {
    "blocked_tools": [
      "delete_*"
    ],
    "approval_tools": []
  }
}

With the Python package installed, you can instead generate a broader starter policy. The command writes the file itself; do not redirect its printed instructions into the JSON file.

Policy CLI
if-policy init -o starter-policy.json
if-policy validate policy.json
if-policy show policy.json

Python quickstart

First create policy.json in your working directory and check the runtime requirements. This walkthrough verifies one allowed call and one blocked call.

Install and configure

Terminal · POSIX shell
python -m venv .venv
source .venv/bin/activate
pip install "inferencefort[langchain]"
unset INFERENCEFORT_KEY
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
export INFERENCEFORT_FAIL_MODE=fail_closed

Create the example

quickstart.py
import inferencefort
from inferencefort import PolicyViolation
from langchain_core.language_models.fake_chat_models import FakeListChatModel

class CountingModel(FakeListChatModel):
    calls: int = 0

    def _call(self, *args, **kwargs):
        self.calls += 1
        return super()._call(*args, **kwargs)

model = CountingModel(responses=["Allowed response"])
print(model.invoke("Say hello").content)
try:
    model.invoke("Synthetic SSN: 123-45-6789")
except PolicyViolation:
    print("Blocked before model execution")
else:
    raise RuntimeError("Expected a policy block")
assert model.calls == 1, "The blocked prompt reached the model"
status = inferencefort.governance_status()
assert status.get("initialised"), "Governance did not initialise"
print("Verified: one allowed call, one blocked call")

Run and verify

Run
python quickstart.py

Expected output includes Allowed response, Blocked before model execution, and the verification message. The counter must remain one. Replace the stand-in with your supported provider only after this succeeds; install that provider integration and configure its endpoint and credentials separately.

TypeScript quickstart

First create policy.json in your working directory and check the runtime requirements. This walkthrough verifies one allowed call and one blocked call.

This executable JavaScript example uses the TypeScript SDK in ESM mode. Explicit wrapModel avoids import-order ambiguity. The model stand-in uses the same model-call surface intercepted in Vercel AI SDK integrations.

Install and configure

Terminal · POSIX shell
npm init -y
npm install @inferencefort/ai-sdk
unset INFERENCEFORT_KEY
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
export INFERENCEFORT_FAIL_MODE=fail_closed

Create the example

quickstart.mjs
import { wrapModel, PolicyViolation, governanceStatus } from '@inferencefort/ai-sdk';

let calls = 0;
const model = wrapModel({
  specificationVersion: 'v2', provider: 'example.chat', modelId: 'example-model',
  async doGenerate() {
    calls++;
    return { content: [{ type: 'text', text: 'Allowed response' }],
      usage: { inputTokens: 1, outputTokens: 1 }, finishReason: 'stop', warnings: [] };
  },
});
const request = text => ({ prompt: [{ role: 'user', content: [{ type: 'text', text }] }] });
await model.doGenerate(request('Say hello'));
try {
  await model.doGenerate(request('Synthetic SSN: 123-45-6789'));
  throw new Error('Expected a policy block');
} catch (error) {
  if (!(error instanceof PolicyViolation)) throw error;
  console.log('Blocked before model execution');
}
if (calls !== 1) throw new Error('The blocked prompt reached the model');
if (!governanceStatus().initialised) throw new Error('Governance did not initialise');
console.log('Verified: one allowed call, one blocked call');

Run and verify

Run
node --import @inferencefort/ai-sdk/register.mjs quickstart.mjs

Use a real Vercel provider

Install ai and @ai-sdk/openai, configure your provider key through your secret manager, and use this model with generateText. This makes a real provider request and may incur usage charges.

Real-provider example

Real-provider example
import { generateText } from 'ai';
import { openai } from '@ai-sdk/openai';
import { wrapModel, withContext, PolicyViolation } from '@inferencefort/ai-sdk';

try {
  const result = await withContext(
    { user: 'developer@example.com', agentId: 'example-agent', threadId: 'example-conversation' },
    () => generateText({ model: wrapModel(openai('gpt-4o')), prompt: 'Say hello' }),
  );
  console.log(result.text);
} catch (error) {
  if (!(error instanceof PolicyViolation)) throw error;
  console.error('Blocked:', error.message);
}

Integration coverage

Python installs lazy interception through the packaged startup hook. Explicitly importing inferencefort or using if-run app.py also activates the hooks. Export configuration before starting the process. Python started without normal site initialization does not load the startup hook.

For Node applications that need automatic interception, load the registration hook before application imports:

Node preload
node --import @inferencefort/ai-sdk/register.mjs app.mjs

Calling activate() after strict ESM imports is not a substitute for the preload hook. Alternatively, use explicit wrappers for the surfaces your application calls.

SurfacePythonTypeScript / NodeVerify
LangChain / LangGraph modelsLazy framework interceptionPreload or explicit framework wrappersInvoke and streaming paths used by your app
Vercel AI SDKNot applicablePreload or wrapModelModel calls and tool execution separately
LiteLLMCompletion and async completionNot applicableActual provider routing
Raw OpenAI / AnthropicSupported client interceptionPreload or wrapOpenAI / wrapAnthropicThe specific API method, including streaming, used by your app
CrewAI toolsSupported tool interceptionNot applicableAgent-loop tools as well as direct calls
MCPClient tool calls and server trace continuationPreload or MCP module wrappersClient/server exchange below
Custom HTTP clients / custom tool loopsDo not assume interception. A supported model client does not prove a custom tool loop is governed.Deliberate blocked action and coverage diagnostics

Other model providers can be reached through supported framework integrations. This does not imply interception of every provider’s raw SDK or every new API method. Recheck your exact paths when upgrading frameworks. Streaming input is evaluated before provider execution; output controls cannot undo effects or content already delivered.

Verify enforcement

  1. Run the local quickstart for your language.
  2. Configure the actual framework, provider, and tool path your application uses.
  3. Run one permitted call and one deliberately prohibited call.
  4. Confirm the prohibited call never reaches its provider or tool handler.
  5. Review status, findings, and audit events. A successful request or an empty findings list alone does not prove coverage.

MCP tools

Complete the Python quickstart first to create the environment and export the policy configuration. For the Node client, also install the SDK from the TypeScript quickstart.

Keep the example policy active. This local server returns fixture text and never deletes anything. Both clients below must print a status result and a block; neither should receive Unexpected execution. Run each client from the directory containing the server file. Stdio transports may filter inherited environment variables, so these examples explicitly pass the absolute policy path to the server without forwarding unrelated credentials.

Python 3.10+ · MCP dependency
pip install "inferencefort[mcp]"
mcp_server.py
import inferencefort
from mcp.server.fastmcp import FastMCP

server = FastMCP("documentation-example")

@server.tool()
def lookup_status() -> str:
    """Return an example status."""
    return "Example status: ready"

@server.tool()
def delete_record() -> str:
    """Fixture only; never deletes data."""
    return "Unexpected execution"

if __name__ == "__main__":
    server.run(transport="stdio")
mcp_client.py
import asyncio
import os
import sys
import inferencefort
from inferencefort import PolicyViolation
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

async def main():
    params = StdioServerParameters(
        command=sys.executable, args=["mcp_server.py"],
        env={"INFERENCEFORT_POLICY_FILE": os.environ["INFERENCEFORT_POLICY_FILE"],
             "INFERENCEFORT_FAIL_MODE": "fail_closed"},
    )
    async with stdio_client(params) as (reader, writer):
        async with ClientSession(reader, writer) as client:
            await client.initialize()
            print(await client.call_tool("lookup_status", {}))
            try:
                await client.call_tool("delete_record", {})
            except PolicyViolation:
                print("Blocked delete_record before transport")
            else:
                raise RuntimeError("Expected a blocked tool call")

asyncio.run(main())
Run Python client
python mcp_client.py

TypeScript SDK client against the same server

Keep the Python environment active so python resolves to the interpreter with MCP installed. Wrappers accept module namespaces, not constructed clients, and return a boolean indicating wrapping success.

Node MCP dependency
npm install @modelcontextprotocol/sdk
mcp_client.mjs
import * as mcpClient from '@modelcontextprotocol/sdk/client/index.js';
import { StdioClientTransport } from '@modelcontextprotocol/sdk/client/stdio.js';
import { wrapMcpClient, PolicyViolation } from '@inferencefort/ai-sdk';

if (!wrapMcpClient(mcpClient)) throw new Error('MCP client wrapping failed');
const client = new mcpClient.Client({ name: 'documentation-client', version: '1.0.0' });
try {
  await client.connect(new StdioClientTransport({ command: 'python', args: ['mcp_server.py'],
    env: { INFERENCEFORT_POLICY_FILE: process.env.INFERENCEFORT_POLICY_FILE,
           INFERENCEFORT_FAIL_MODE: 'fail_closed' } }));
  console.log(await client.callTool({ name: 'lookup_status', arguments: {} }));
  try {
    await client.callTool({ name: 'delete_record', arguments: {} });
    throw new Error('Expected a blocked tool call');
  } catch (error) {
    if (!(error instanceof PolicyViolation)) throw error;
    console.log('Blocked delete_record before transport');
  }
} finally {
  await client.close();
}
Run Node client
node --import @inferencefort/ai-sdk/register.mjs mcp_client.mjs

Instrument a Node MCP server

Wrap the low-level server module before constructing the server and registering its handlers. Continue your normal MCP server setup after this fragment. Inbound trace continuation does not replace authentication or server authorization.

Node server initialization · fragment
import * as mcpServer from '@modelcontextprotocol/sdk/server/index.js';
import { wrapMcpServer } from '@inferencefort/ai-sdk';

if (!wrapMcpServer(mcpServer)) throw new Error('MCP server wrapping failed');
const server = new mcpServer.Server(
  { name: 'example-server', version: '1.0.0' },
  { capabilities: { tools: {} } },
);
// Register your tool handlers and connect your transport here.

Controls

Model access, content rules, configured-destination restrictions, and tool policies can be evaluated locally from policy. Configured detectors, session services, and shared budgets may require network requests. The controls available depend on your account, deployment, and integration coverage.

ControlConfiguration and behavior
Model accessAllow or deny models by policy. Include an explicit allow policy for intended usage; unmatched access is denied when policy is loaded.
Content rulesSubstring or regex rules with block or flag actions, scoped to input, output, or both. Output decisions occur after execution.
EgressProvider-to-destination allow-lists check the configured destination. Matching is host-anchored; loopback is allowed by default. Pair with network controls.
MCP serversAn empty allow-list is unrestricted; stdio/in-process calls are local. Configure remote servers explicitly.
Tool policyblocked_tools and approval_tools accept tool-name patterns. An approval-required verdict stops the call; build the human review and authorized retry into your application.
PHI and detectorsEnable the intended scanners and redaction separately. Install/configure the required redactor or detector endpoint and check capability gaps. Detection does not guarantee every sensitive value is found.
Session detectionAccount-dependent controls use session context. Findings and configured actions determine blocking, flagging, or approval; do not assume every finding is removed from returned text.
BudgetsShared daily caps need the control plane. Local-only operation does not provide a shared ledger.

Policy fragments

Merge the relevant fragment into your reviewed policy; these are not standalone complete policies.

Destinations and tools · fragment
{
  "egress_endpoints": {"openai": ["https://api.openai.com/v1"]},
  "mcp_servers": ["https://tools.example.com"],
  "mcp_policy": {"blocked_tools": ["delete_*"], "approval_tools": ["send_email"]}
}

Use enforcement_mode: "monitor" to evaluate without refusing policy violations, and "enforce" to block. Review monitor findings before enabling enforcement. Do not use monitor mode for the blocked-call walkthroughs.

For PHI in Python, install inferencefort[hipaa] and the selected spaCy model, then enable redaction in policy. Node deployments must configure their supported redaction service when required, including INFERENCEFORT_PHI_URL. Confirm capability status rather than assuming equivalent redaction coverage from package installation.

Connect a control plane

Use the control-plane address supplied for your hosted workspace or your own deployment. The example address below is a placeholder, not a working service. Set keys through your deployment secret manager; never commit them into application code, policy files, or documentation.

Hosted or self-hosted configuration
export INFERENCEFORT_API_URL="https://control-plane.example.com"
# Set INFERENCEFORT_KEY through your secret manager before starting the app.

The default API address is http://localhost:8080; setting a key does not change it. A provider key is separate from your InferenceFort key. Verify connectivity, policy availability, and the workspace’s enforcement mode before relying on a blocked action.

Backend policy refreshes on a server-controlled TTL. Changes are not instantaneous. Existing cached rules continue to apply during refresh failures; verify the active policy before judging a change. Audit delivery is asynchronous and should not be treated as proof of durable receipt merely because a call returned.

Local policy plus backend policy

A configured local file can remain active alongside backend policy; either layer can tighten restrictions. For an agent fleet, set INFERENCEFORT_POLICY_DIR to a directory of JSON files declaring agent_id, and optionally system_id to disambiguate systems. Do not put backend-managed pricing, budget, detector, classification, or kill-switch settings into new per-agent files. Validate each file before deployment.

Identity & isolation

Set identity from your authenticated server context before the agent runs. Do not trust a tenant or user identifier supplied in an arbitrary request body. Session state is keyed by customer and conversation, so reuse the conversation ID across requests in the same conversation and separate customers explicitly.

Request scoping

Python request boundary · application fragment
import inferencefort

async def run_agent(authenticated_user, authorized_customer, conversation_id, message):
    with inferencefort.context_scope(
        user=authenticated_user,
        customer_id=authorized_customer,
        agent_id="support-example",
        thread_id=conversation_id,
    ):
        return await agent.ainvoke(message)  # your configured agent
Node request boundary · application fragment
import { withContext } from '@inferencefort/ai-sdk';

async function runAgent(authenticatedUser, authorizedCustomer, conversationId, message) {
  return withContext(
    { user: authenticatedUser, customerId: authorizedCustomer,
      agentId: 'support-example', threadId: conversationId },
    () => agent.invoke(message), // your configured agent
  );
}

Identity is optional unless policy requires it. Missing an explicit conversation ID causes a fallback ID to be pinned to the current context; later calls in that context reuse it. The fallback does not automatically join separate HTTP requests into one conversation.

Python contextvars propagate through await, into tasks created under the context, and through asyncio.to_thread. Raw executor submissions need explicit propagation. Set identity before spawning work, not after the call or in a sibling task.

Python raw executor · fragment
from inferencefort.core.context import bind_context

await loop.run_in_executor(pool, bind_context(lambda: agent.invoke(message)))

Keep request scope outside the graph run so child nodes inherit it. For LangGraph, pass invocation configuration explicitly as a keyword in Python. For Node worker threads or separate processes, transfer correlation context explicitly rather than expecting AsyncLocalStorage to cross that boundary.

Cross-agent traces

MCP integrations can propagate trace context through request metadata. For your own HTTP or queue transport, inject the context on the caller and extract it on the receiver. These fragments assume authenticated transport and an existing governed caller scope.

Python correlation · application fragment
from inferencefort import inject_context, extract_context, remote_context

headers = {}
inject_context(headers)
# Send headers with your authenticated request.

# Receiver: establish its own authorized customer identity first.
with remote_context(extract_context(incoming_headers)):
    result = agent.invoke(message)
Node correlation · application fragment
import { injectContext, extractContext, withRemoteContext } from '@inferencefort/ai-sdk';

const headers = {};
injectContext(headers);
// Send headers with your authenticated request.

// Receiver: establish its own authorized customer identity first.
const result = await withRemoteContext(
  extractContext(incomingHeaders), () => agent.invoke(message),
);

Correlation headers are not authentication. The receiver must establish its own customer identity. Cross-process session continuation depends on compatible trust configuration; trace linkage alone does not prove session state is shared. Coordinate shared session configuration across participating deployments and verify an end-to-end trace.

Offline deployment

No-key local policy supports static model access, content rules, egress, MCP server restrictions, and blocked/approval tools. It does not enable account-dependent session detection, including the lethal-trifecta guard, value provenance, runtime tool classification, or session-risk classification. It also does not provide shared budgets or the hosted audit trail.

Configured customer detectors, redaction services, and SIEM forwarding may still make network requests. “No InferenceFort account” and “no outbound network” are different deployment properties. For air-gapped use, keep required services within the permitted boundary, review every configured endpoint, and verify the supplied deployment’s supported capabilities.

Restart after changing a static policy file. To propose policy from observed audit traffic, use if-policy learn audit.ndjson; review the proposal and validate it before switching from monitor to enforcement. Audit files can contain sensitive content and must follow your organization’s access and retention rules.

Troubleshooting

Inspect status after a deliberate governed call: initialization may be lazy. Check both initialization and capability gaps. Initialization confirms a loaded policy, not coverage of every framework or tool path.

Python diagnostics
import inferencefort
status = inferencefort.governance_status()
print("Initialised:", status.get("initialised", False))
print("Capability gaps:", status.get("capability_gaps", {}))
TypeScript SDK diagnostics
import { governanceStatus } from '@inferencefort/ai-sdk';
const status = governanceStatus();
console.log('Initialised:', status.initialised);
console.log('Capability gaps:', status.capabilityGaps);
SymptomCheck
Installed, but no blockConfiguration before startup, loaded policy, enforce mode, matching rule, and the actual intercepted call path.
Not initializedPolicy path or control-plane connectivity, native library availability, and release compatibility with this runtime.
ESM model calls pass throughPreload registration before imports or explicitly wrap the model.
Model governed, tool ungovernedConfirm the tool integration separately; model interception does not wrap arbitrary application actions.
Rule misses expected textCheck input/output scope and scanned_scopes. System prompts and assistant history require the relevant scanning configuration.
Redaction or detector unavailableReview capability gaps, required dependencies, endpoint reachability, and enabled configuration.
Wrong user or sessionSet authenticated identity inside the request scope before creating child tasks; verify trace attribution.
No hosted audit eventsCheck account configuration, connectivity, content-logging settings, and asynchronous delivery. No-key mode has no hosted audit trail.

For focused troubleshooting, enable stage logging temporarily. Restrict access to diagnostic output, inspect it for sensitive content before sharing, and disable it when finished.

Optional local diagnostics
export INFERENCEFORT_STAGE_LOG=1
export INFERENCEFORT_STAGE_LOG_FILE=./stages.ndjson

Fail modes

Choose one mode for your deployment. Cached local rules are still evaluated; a remote outage does not automatically disable all enforcement. Configured detectors and other services can have their own failure settings, which also need review.

ModeImportant behavior
fail_open (default)Availability-oriented fallback when governance cannot obtain a required answer. Calls may proceed without the unavailable control.
fail_closedRefuses calls when a required policy cannot be loaded or a budget check fails. Test the resulting application errors before enabling.
fail_cachedFor budget outages, reuses the last successful budget answer when available. Without a cached budget answer, the call can proceed. It is not a universal replay of every previous verdict.

Catch PolicyViolation to handle a refused action. Rethrow unrelated exceptions; provider and application failures remain distinct. For approval-required actions, present the proposed action for authorized review and implement the supported approval/retry flow. Never treat merely catching an exception as approval.

Environment variables

VariablePurpose
INFERENCEFORT_KEYWorkspace SDK credential; inject through a secret manager.
INFERENCEFORT_API_URLControl-plane base URL. Default http://localhost:8080.
INFERENCEFORT_POLICY_FILEPath to a reviewed static policy bundle.
INFERENCEFORT_POLICY_DIRDirectory of per-agent JSON policies.
INFERENCEFORT_FAIL_MODEfail_open, fail_closed, or fail_cached.
INFERENCEFORT_CUSTOMER_IDDeployment-level customer identity where appropriate; use request scope for multiple customers.
INFERENCEFORT_PHI_URLNode delegated redaction endpoint when required by your deployment.
INFERENCEFORT_MCP_PROPAGATESet to 0 to disable automatic MCP correlation propagation.
INFERENCEFORT_SHARED_TAINTShared session capability configuration; coordinate across participating agents.
INFERENCEFORT_THREAD_MAXBound on retained sessions; default 10000.
INFERENCEFORT_TRACE_BUFFEROpt in to local trace buffering; off by default.
INFERENCEFORT_STAGE_LOG / INFERENCEFORT_STAGE_LOG_FILEEnable and choose the destination for local stage diagnostics.

This is the configuration used in these guides, not a promise that every framework or release accepts every optional setting. Keep the SDK, native library, and framework versions aligned with your supplied release.

© 2026 InferenceFort About