AI Security for On-Prem Models
Local models still follow injected instructions.
Running Llama, Qwen or Gemma on your own hardware keeps prompts inside your network. It does not stop an agent from reading a poisoned document and calling a tool it should not. InferenceFort governs agents on local models the same way it governs hosted ones.
What running on-prem does and does not solve
Your prompts stay home
Prompts are not sent to a model vendor, vendor retention is out of the picture, and third-party inference egress disappears.
Your agent can still be turned
Prompt injection, unsafe tool calls, unapproved MCP servers, data sent onward by the agent's own tools, and the lack of an audit trail.
Govern local endpoints like any other
Ollama, vLLM and similar servers are usually reached through OpenAI-compatible clients or framework integrations such as LangChain's ChatOllama. Those calls are intercepted exactly like calls to hosted providers. Loopback endpoints such as localhost and 127.0.0.1 are always allowed by egress policy because they cannot leave the machine; an internal host such as gpu-01.internal is allowed when you list it.
Keep detection inside your network
- Local prompt-attack detector: point it at your own OpenAI-compatible model, including
http://localhost:11434/v1. - Your own classifier: bring an OpenAI-compatible model for the session-risk classifier, configured per workspace.
- In-process detectors: register a function that calls a local model, a library or your own rules. It runs in your process and is never fetched over the wire.
- Internal detection APIs: describe an HTTP detector in configuration. A descriptor can name a destination and select fields; it cannot carry code.
- Learned tool profiles: generated by a local model by default. Pointing generation at a hosted model is an explicit opt-in.
Choose a deployment mode
| Connected | Offline (policy file) | |
|---|---|---|
| Network | SDK talks to a hosted or self-hosted control plane for policy, budgets and audit | Zero network calls to InferenceFort |
| Static policy | Model access, content rules, egress, MCP allow-list, tool blocking | Same |
| Detectors and PHI redaction | Yes, on the endpoints you configure | Yes, on the endpoints you configure |
| Session-based detection | Lethal-trifecta guard, tool classification, session-risk classifier | Not available |
| Budgets, live policy changes, hosted trail | Yes | Not available; use a file or SIEM sink as the record |
Offline mode never pretends. Stages that are withheld are marked as skipped with a reason and reported as a capability gap, and they never cause blocks, even under fail_closed.
if-policy init -o policy.json
if-policy validate policy.json
export INFERENCEFORT_POLICY_FILE="$PWD/policy.json"
unset INFERENCEFORT_KEY"siem": {"provider": "file", "enabled": true, "mode": "siem_only",
"vars": {"path": "./audit.ndjson"}}Frequently asked questions
Does it work in an air-gapped environment?
Yes, in offline mode: static policy, detectors on your endpoints, PHI redaction and a local audit sink, with no network calls to InferenceFort.
In connected mode, what leaves our network?
Policy fetches, budget checks and audit events go to your control plane, which can be self-hosted. Detectors call only the endpoints you configure.
Which local models are supported?
Any model reached through a supported framework or client, such as LangChain, LiteLLM, the Vercel AI SDK or the OpenAI SDK pointed at an OpenAI-compatible server.
Keep reading
Plan an on-prem evaluation
Bring one agent workflow, the data it touches, and the actions you need to control. As a design partner, you shape the evaluation and review the policy decisions with our engineers.
Become a design partner →