We shipped a connector for Baseten. If your models run there — a fine-tuned Llama on a dedicated deployment, DeepSeek through Model APIs, a multi-model Chain — Wingback now discovers them, guards them in-flight, and red-teams them like everything else in your AI estate.
The model you trained is the model nobody watches
The center of gravity of enterprise inference has been moving. Frontier APIs still matter, but a growing share of production traffic goes to open-weight and fine-tuned models on dedicated inference platforms, where the economics and the latency are better and the weights are yours.
Security tooling has not moved with it. Gateways, guardrails, and monitoring all grew up pointed at the big provider APIs. The result is a quiet inversion: the model you rent from a frontier lab arrives with that lab’s safety training and sits behind everyone’s tooling, while the model you host yourself — the one tuned on your data, serving your customers — often runs with neither. A model you host is not a model you can ignore. It is the one you most need to watch.
What the connector does
The connector plugs Baseten into all four layers of the platform:
- Discovery. It walks your Baseten workspace and lands every dedicated deployment, Model API endpoint, and Chain in the same inventory as your agents, models, and MCP connectors. Deployments nobody wrote down surface next to the ones everybody knows about.
- In-flight guarding. Baseten endpoints speak the OpenAI API shape (with Anthropic-compatible endpoints in beta), so the Wingback gateway drops in front of them the way it fronts any provider — a base-URL change, not an integration project. Prompt injection, tool misuse, and data-egress guards now apply to the models you host.
- Adaptive red team. The engine that attacks your frontier-API agents now attacks your Baseten-served models and Chains. Open and fine-tuned models fail differently from lab-hardened frontier models — tuning can erode refusal behavior in ways nobody notices until someone looks. Every validated finding compiles into a runtime guard, so the same attack cannot run twice.
- Compliance mapping. Findings from Baseten-served models map to the OWASP LLM and Agentic Top 10 and MITRE ATLAS alongside everything else, in the same reports.
Why Baseten first
Baseten is where a serious share of open-model inference now runs, and its API surface made the integration clean: OpenAI-compatible endpoints meant our gateway and red-team engine pointed at it almost unchanged. Chains matter too — multi-model workflows are agent-shaped, and agent-shaped things are exactly what we built the platform to watch.
We wrote last month about what concentrates inside an LLM gateway, and that the biggest AI security failures live in the infrastructure around the models. This is the same thesis pointed forward: coverage should follow the model wherever it serves, not stop at the API bill from a frontier lab.
If you run models on Baseten, we can show you what Wingback sees in your workspace — request a demo. If you would rather read first, start with the one-pager.
Baseten platform details: Model APIs and the Baseten inference stack. Baseten is a trademark of Baseten, Inc.; this connector is built and supported by Wingback Security.

