LLM Runtime Security

Enforce policy at the moment your application calls the model.

Gateways see traffic after it leaves. Prompt filters see text without context. InferenceFort runs inside your application, where it sees the real call: which model, which endpoint, which user, which tool. It decides before the request is sent.

Why enforce in process

Nothing to route around

A gateway only governs traffic sent through it. A direct provider call or a framework's own client skips it. Interception at the framework layer governs the call wherever it goes.

Full context

In process, InferenceFort knows the user, agent, customer and conversation, and what the agent did earlier in the session. A network proxy sees a request body.

No round trip

Policy arrives as a cached bundle and is evaluated locally. Allowed calls do not wait on a network hop.

One engine, every language

Decisions are implemented once, in a dependency-free Go core loaded by both the Python and TypeScript SDKs. The same policy gives the same answer in every service.

What happens on each call

The governed call path
  1. Your app calls a model or tool through a supported framework or client. Python patches frameworks as they are imported; Node uses a register hook or explicit wrappers.
  2. The SDK resolves the provider, the configured destination and the caller's identity.
  3. The core evaluates model access, content rules, egress, detectors and tool policy from the cached bundle.
  4. A blocked call raises PolicyViolation, the only exception InferenceFort raises into your code. The request never leaves.
  5. After the call, usage and the decision are recorded asynchronously. Cost records survive outages through a bounded local write-ahead log.

Controls enforced at runtime

ControlWhat it does
Model accessAllow or deny models per user, agent or team. When policy is loaded, unmatched access is denied.
Content rulesSubstring and regex rules that block or flag, on input, output or both.
EgressPin each provider to approved endpoints or regions. Host-anchored matching; local loopback is always allowed.
DetectorsBuilt-in local prompt-attack detector, Lakera, HTTP detection APIs described in config, or your own in-process function.
PHI redactionShaped identifiers are replaced with typed placeholders such as [SSN] before the model sees the prompt.
Tool policyBlocked and approval-required tools, MCP server allow-lists and session-based exfiltration checks.
BudgetsDaily spend caps checked against a shared ledger. The only per-call network check.

Coverage

SurfacePythonTypeScript / Node
LangChain / LangGraphEvery BaseChatModel provider, including streamingLangChain.js and LangGraph.js, including streaming
Vercel AI SDK–doGenerate and doStream
LiteLLMcompletion / acompletion, plus a proxy guardrail–
OpenAI / Anthropic SDKsSupported client methodsRegister hook or wrapOpenAI / wrapAnthropic
ToolsLangChain, CrewAI, MCPMCP

Get started

Python
pip install "inferencefort[langchain]"
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"   # or set INFERENCEFORT_KEY
python app.py                                          # or: if-run app.py
Node
npm install @inferencefort/ai-sdk
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
node --import @inferencefort/ai-sdk/register.mjs app.mjs

Built to stay out of the way

Frequently asked questions

Do I have to pass callbacks to every model?

No. Supported frameworks are intercepted at their shared entry points, so every provider built on them is governed without per-model code.

Is streaming governed?

Yes, on the supported streaming paths. Input is checked before the provider runs; output rules run as the response is produced.

What if my framework is not listed?

Many providers are reachable through a supported framework such as LangChain, LiteLLM or the Vercel AI SDK. A custom HTTP client or custom tool loop is not governed automatically; verify coverage with a deliberate blocked call.

Keep reading

Validate runtime coverage

Bring one agent workflow, the data it touches, and the actions you need to control. As a design partner, you shape the evaluation and review the policy decisions with our engineers.

Become a design partner →