Red Hat has announced significant updates across the Red Hat AI portfolio with the release of Red Hat AI 3.5. As enterprise teams move past early experimentation and pilot successes, IT and platform engineering leaders face the challenge of running AI with the same operational rigor as mission-critical infrastructure. By providing the scalable foundation required to control, secure and observe these workloads across the hybrid cloud, Red Hat AI 3.5 bridges the gap between isolated AI pilots and a fully governed enterprise architecture.
Red Hat AI 3.5 delivers the operational foundation organizations need to scale AI in production and extend it across hybrid environments through new safety and observability capabilities. With this release, organizations can verify models before deployment through EvalHub, enabling risk-focused safety benchmarking and the creation of regulatory compliance certifications. New observability dashboards give platform teams comprehensive metrics to gain real-time insight into inference health, GPU utilization, and AI model performance. Non-admin users can access dashboards for per-user token consumption showback, and distributed inference workloads.
In addition, Red Hat AI 3.5 expands on proven enterprise platform capabilities to deliver enhanced multi-tenancy for AI service providers and AI use cases that require complete hardware-to-software isolation as well as priority-aware serving with native multi-tenancy for shared GPU infrastructure. For organizations that require stronger isolation between tenants, Red Hat AI now officially supports running on Red Hat OpenShift hosted control planes deployed on Red Hat OpenShift Virtualization. Hosted control planes give every tenant a dedicated cluster control plane while consolidating the hardware beneath them. Running AI workloads in Red Hat OpenShift Virtualization virtual machines adds robust VM-level isolation across shared, GPU-enabled infrastructure. Together, these capabilities let infrastructure providers operate and upgrade the entire underlying environment from a single point of control.
Red Hat AI 3.5 also accelerates the path to governed AI agents. AutoRAG links enterprise data repositories straight to agentic applications, introducing advanced capabilities such as multilingual document support, conversational testing, and contextual retrieval. A visual pipeline gives teams confidence in their RAG configurations before deploying. Agent templates deliver pre-configured implementations for common patterns like code review, document processing, and research workflows. Once deployed, Inference-Time Scaling optimizes GPU spend by adjusting compute dynamically based on query complexity.
Why does Red Hat AI 3.5 matter?
As enterprise AI pilots succeed and initial results show promising returns, IT teams must then address the need for delivery at scale. But scaling AI across the business demands the same operational rigor as any mission-critical infrastructure: verified safety before deployment, precise resource controls across shared GPU environments, governed agent behavior and transparent usage metrics. Most organizations today face these challenges with fragmented tooling and manual processes that don't scale. Red Hat AI 3.5 addresses this gap by unifying safety, multi-tenancy, agentic development and observability into a single enterprise platform. The result is AI that operates as an accountable, governed AI architecture — not an isolated experiment — across hybrid cloud environments.
“The conversation has moved from getting AI into production to running it at scale as trusted enterprise infrastructure, which requires safety evidence, governed agents, cost attribution and multi-tenancy,” said Joe Fernandes, vice president and general manager, AI Business Unit, Red Hat. “With Red Hat AI 3.5, we are delivering the operational controls, verifiable trust and agentic foundations IT leaders need to run AI as a safe, controlled and accountable enterprise AI architecture across the hybrid cloud.”
Key takeaways
• Verifiable pre-deployment safety and evaluation: Evaluated catalog models feature built-in Garak benchmark scores, while the general availability of EvalHub automates safety and auditable compliance reporting for custom models, RAG and agents.
• Shared GPU control for multi-tenant inference: Fair-share GPU scheduling manages resource allocation across tenants, while priority-aware serving provides admission control and priority-based request routing to protect real-time inference and allows background workloads to use available capacity.
• Agent APIs and gateway security: General availability support for the Responses API and built-in RAG provides a unified open-source interface for multi-turn agent conversations, reinforced by integrated NeMo Guardrails that intercept malicious tool calls.
• Enterprise data grounding and efficient reasoning: AutoRAG with pgvector support, native AutoML, and Inference-Time Scaling (ITS) allow models to adapt compute usage dynamically based on query difficulty.
• Built-in observability and MaaS showback: Delivers per-user token metering, performance dashboards for models and agents, MLflow visual agentic tracing, and GPU utilization dashboards for clear operational and usage transparency.
• Pre-built agent templates for faster development: AI Hub introduces agent templates and starter kits with pre-configured reference implementations for common enterprise patterns, including code review, document processing, and research workflows.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.




