Singapore – New Relic has launched AI Evaluation, a framework that scores AI response quality and guardrail performance across entire application transactions, the company announced on 7 October 2026.
The capability sits within New Relic AI Observability and tracks response quality and behaviour from development through to production. New Relic says it also provides automated, real-time insights into how AI guardrails perform.
The company positions the tool against point solutions that evaluate isolated single-LLM calls. Instead, it says, AI Evaluation links results to business impact across the full developer-to-production lifecycle.
New Relic points to a wider shift behind the launch. As generative AI becomes embedded in enterprise workflows, software is moving from deterministic code to probabilistic systems, which can fail in ways ranging from subtle hallucinations to unexpected performance shifts when frontier vendors update their LLMs.
As a result, traditional application signals no longer tell the whole story, the company says. Teams now need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost and whether the result was useful.
New Relic argues that observing AI in a silo, or through best-of-breed point solutions, creates blind spots around upstream and downstream operational impacts. It adds that this approach widens the visibility gap between AI developers and production engineering.
AI Evaluation is built natively into the New Relic platform and analyses AI performance down to the underlying transaction. According to the company, it embeds quality and security guardrails directly into existing application performance monitoring workflows.
The system replaces manual reviews with an asynchronous “LLM-as-a-judge” service. It scans live telemetry to score vulnerabilities such as hallucinations, prompt injections and data leaks.

These probabilistic quality scores attach as attributes to deterministic distributed traces. New Relic says this allows teams to isolate the root cause of a failure, whether in the prompt, the vector database or backend infrastructure, within a single view.
The platform also links qualitative response scores to compute consumption. In turn, New Relic says, teams can judge whether expensive models deliver enough semantic value over faster, lower-cost alternatives.
“Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime,” said New Relic Chief Product Officer Brian Emerson.
“Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact, and performance—all within the platform tools that SREs, platform engineers, and developers already use,” Emerson said.
In production, configurable guardrails evaluate sampled inputs and outputs to help spot malicious prompt injections, jailbreak attempts, accidental PII leaks, toxicity and bias. New Relic says this lets engineering teams act before such issues cause reputational damage or regulatory fines.
For retrieval-augmented generation (RAG) architectures, metrics such as faithfulness and answer relevancy help engineers separate a model’s reasoning performance from a vector database’s retrieval logic. Pre-built evaluators are also included to simplify configuration.
Before deployment, a prompt playground allows engineers to test and refine prompts against real models side by side in a secure environment. Teams can also curate and version “golden datasets”, sourced from real distributed traces or synthetic data, to run regression tests on any changes.
Additionally, organisations can run controlled A/B tests across different prompts, models and configurations. New Relic says this reduces the guesswork of prompt engineering.
Industry analysts have also commented on the shift. “As organisations move generative AI applications from early pilots into mission-critical production environments, traditional application performance metrics are no longer sufficient on their own,” said Stephen Elliot, Group Vice President, I&O, Cloud Operations, and DevOps of IDC.
“A successful AI implementation requires visibility into both technical health and response quality, including accuracy, safety, and model efficiency. Bridging live response evaluation with prompt lifecycle management and full-stack operational telemetry is becoming essential for enterprise engineering and security teams looking to mitigate risk and manage costs effectively,” Elliot concluded.

