AI Security for On-Prem Models

Local models still follow injected instructions.

Running Llama, Qwen or Gemma on your own hardware keeps prompts inside your network. It does not stop an agent from reading a poisoned document and calling a tool it should not. InferenceFort governs agents on local models the same way it governs hosted ones.

What running on-prem does and does not solve

Solved by self-hosting

Your prompts stay home

Prompts are not sent to a model vendor, vendor retention is out of the picture, and third-party inference egress disappears.

Not solved by self-hosting

Your agent can still be turned

Prompt injection, unsafe tool calls, unapproved MCP servers, data sent onward by the agent's own tools, and the lack of an audit trail.

Govern local endpoints like any other

Ollama, vLLM and similar servers are usually reached through OpenAI-compatible clients or framework integrations such as LangChain's ChatOllama. Those calls are intercepted exactly like calls to hosted providers. Loopback endpoints such as localhost and 127.0.0.1 are always allowed by egress policy because they cannot leave the machine; an internal host such as gpu-01.internal is allowed when you list it.

Keep detection inside your network

Choose a deployment mode

ConnectedOffline (policy file)
NetworkSDK talks to a hosted or self-hosted control plane for policy, budgets and auditZero network calls to InferenceFort
Static policyModel access, content rules, egress, MCP allow-list, tool blockingSame
Detectors and PHI redactionYes, on the endpoints you configureYes, on the endpoints you configure
Session-based detectionLethal-trifecta guard, tool classification, session-risk classifierNot available
Budgets, live policy changes, hosted trailYesNot available; use a file or SIEM sink as the record

Offline mode never pretends. Stages that are withheld are marked as skipped with a reason and reported as a capability gap, and they never cause blocks, even under fail_closed.

Offline setup
if-policy init -o policy.json
if-policy validate policy.json
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
unset INFERENCEFORT_KEY
policy.json · local audit sink fragment
"siem": {"provider": "file", "enabled": true, "mode": "siem_only",
         "vars": {"path": "./audit.ndjson"}}

Frequently asked questions

Does it work in an air-gapped environment?

Yes, in offline mode: static policy, detectors on your endpoints, PHI redaction and a local audit sink, with no network calls to InferenceFort.

In connected mode, what leaves our network?

Policy fetches, budget checks and audit events go to your control plane, which can be self-hosted. Detectors call only the endpoints you configure.

Which local models are supported?

Any model reached through a supported framework or client, such as LangChain, LiteLLM, the Vercel AI SDK or the OpenAI SDK pointed at an OpenAI-compatible server.

Keep reading

Plan an on-prem evaluation

Bring one agent workflow, the data it touches, and the actions you need to control. As a design partner, you shape the evaluation and review the policy decisions with our engineers.

Become a design partner →