New Relic adds AI evaluation to observability platform to track transaction-level impact of AI workloads

by

Bernard Parado

-

15 hours ago

Singapore – New Relic has launched AI Evaluation, a framework that scores AI response quality and guardrail performance across entire application transactions, the company announced on 7 October 2026.

The capability sits within New Relic AI Observability and tracks response quality and behaviour from development through to production. New Relic says it also provides automated, real-time insights into how AI guardrails perform.

The company positions the tool against point solutions that evaluate isolated single-LLM calls. Instead, it says, AI Evaluation links results to business impact across the full developer-to-production lifecycle.

New Relic points to a wider shift behind the launch. As generative AI becomes embedded in enterprise workflows, software is moving from deterministic code to probabilistic systems, which can fail in ways ranging from subtle hallucinations to unexpected performance shifts when frontier vendors update their LLMs.

As a result, traditional application signals no longer tell the whole story, the company says. Teams now need to see which model responded, which prompt was used, what tools an agent called, how much the interaction cost and whether the result was useful.

New Relic argues that observing AI in a silo, or through best-of-breed point solutions, creates blind spots around upstream and downstream operational impacts. It adds that this approach widens the visibility gap between AI developers and production engineering.

AI Evaluation is built natively into the New Relic platform and analyses AI performance down to the underlying transaction. According to the company, it embeds quality and security guardrails directly into existing application performance monitoring workflows.

The system replaces manual reviews with an asynchronous “LLM-as-a-judge” service. It scans live telemetry to score vulnerabilities such as hallucinations, prompt injections and data leaks.

These probabilistic quality scores attach as attributes to deterministic distributed traces. New Relic says this allows teams to isolate the root cause of a failure, whether in the prompt, the vector database or backend infrastructure, within a single view.

The platform also links qualitative response scores to compute consumption. In turn, New Relic says, teams can judge whether expensive models deliver enough semantic value over faster, lower-cost alternatives.

“Generative AI has redefined application health, making evaluation of response quality and model efficiency just as critical as uptime,” said New Relic Chief Product Officer Brian Emerson.

“Unlike point solutions that evaluate isolated LLM calls, we look holistically across the entire application transaction to show exactly what happens whenever AI is involved, giving teams full visibility into technical health, business impact, and performance—all within the platform tools that SREs, platform engineers, and developers already use,” Emerson said.

In production, configurable guardrails evaluate sampled inputs and outputs to help spot malicious prompt injections, jailbreak attempts, accidental PII leaks, toxicity and bias. New Relic says this lets engineering teams act before such issues cause reputational damage or regulatory fines.

For retrieval-augmented generation (RAG) architectures, metrics such as faithfulness and answer relevancy help engineers separate a model’s reasoning performance from a vector database’s retrieval logic. Pre-built evaluators are also included to simplify configuration.

Before deployment, a prompt playground allows engineers to test and refine prompts against real models side by side in a secure environment. Teams can also curate and version “golden datasets”, sourced from real distributed traces or synthetic data, to run regression tests on any changes.

Additionally, organisations can run controlled A/B tests across different prompts, models and configurations. New Relic says this reduces the guesswork of prompt engineering.

Industry analysts have also commented on the shift. “As organisations move generative AI applications from early pilots into mission-critical production environments, traditional application performance metrics are no longer sufficient on their own,” said Stephen Elliot, Group Vice President, I&O, Cloud Operations, and DevOps of IDC.

“A successful AI implementation requires visibility into both technical health and response quality, including accuracy, safety, and model efficiency. Bridging live response evaluation with prompt lifecycle management and full-stack operational telemetry is becoming essential for enterprise engineering and security teams looking to mitigate risk and manage costs effectively,” Elliot concluded. 

Recognise the innovators redefining commerce at the Retail & E-commerce Excellence Awards Asia Pacific 2026! Taking place this December 2026, we celebrate the region’s most impactful retail strategies, standout e-commerce experiences, and forward-thinking leaders—submit your entries today!
Honour the women shaping the future of marketing and technology at the Empowered Women Awards 2026! This December 2026, we celebrate inspiring leaders, changemakers, and rising voices driving impact across the industry—submit your entries today!
Share

RECENT ARTICLES

Sumsub launches Reusable KYC gateway with MiniPay as first live integrator
Tech Mahindra expands Gemini Enterprise readiness across 12,500 associates
Smartstream takes Air live at Rothera to automate derivatives reconciliations
Securitize and LG CNS to explore tokenised asset infrastructure in South Korea
Global Payments extends cross-border payments to Thailand through Centara Hotels & Resorts deal 
Ellipse 3

RELATED ARTICLES

Everpure launches data management tools to move enterprise AI into production
palantir armada
New Relic names Simon Rizkalla as vice president of customer advocacy for APJ
Ellipse 3

FEATURED ARTICLES

'Retail & E-Commerce Innovation Summit' returns for its 2nd edition in the Philippines — initial speaker lineup revealed
‘Retail & E-Commerce Excellence Awards Asia Pacific
Empowered Women Awards 2026 to honour leading women trailblazing the technology and marketing industries

Subscribe to UpTech Media Newsletter

JOIN OUR NEWSLETTER

Subscribe to our newsletter to get the latest APAC marketing news.