Go to All Forums

Monitor LLM performance, usage, and costs with Site24x7 LLM Observability

Hello, Site24x7 community!

Large language models (LLMs) are increasingly being used to generate content, answer queries, automate workflows, and power AI-driven experiences. As applications begin interacting with multiple LLMs, keeping track of their performance, token consumption, reliability, and costs can become challenging.

To simplify this, we’re introducing LLM Observability in Site24x7, providing centralized visibility into your LLM interactions. Analyze requests across models, track latency and token consumption, and understand model-wise costs. Use LLM Observability to identify errors and investigate individual LLM calls and requests—all from a single console.

What can you do with LLM Observability?


With Site24x7 LLM Observability, you can:

  • Monitor LLM performance: Track request volume and response latency, including latency percentiles, across models to identify slow-performing models and unusual changes in performance.
  • Analyze token consumption: Monitor input and output token usage across individual models and identify workloads with high token consumption.
  • Understand LLM costs: Analyze derived total cost, average cost per request, and model-wise costs to identify models or workloads contributing significantly to your LLM expenditure.
  • Compare model performance: Compare models based on request count, latency, token consumption, and cost from a single view.
  • Identify reliability issues: Monitor errors and identify periods with increased LLM request failures.
  • Forecast token usage: Analyze historical token consumption and forecast future usage to plan LLM resource consumption and associated costs.
  • Investigate individual interactions: Use Prompt Analysis to examine individual LLM requests, including prompt content, model, service, token usage, cost, latency, and request time.
  • Troubleshoot individual LLM interactions: Use Prompt Analysis to drill down from aggregated metrics into individual LLM requests and review prompts, responses, token usage, costs, and latency. Click any record in the Prompt Analysis tab and select the Trace tab to visualize the complete request flow, and use individual spans to pinpoint slow operations and identify performance bottlenecks.



  • Analyze LLM-related logs: Query LLM-related logs to gain additional context while investigating requests and errors.

Use case: Troubleshooting increased LLM latency

Challenge

Suppose users begin experiencing slower responses from an AI-powered application.

Determining whether the slowdown is caused by a particular LLM, a subset of slower requests, or another operation in the application flow can require analyzing multiple layers of data.

Solution

With Site24x7 LLM Observability, you can start with Avg Latency to determine whether overall request latency has increased. Use Response Time by Model to identify the affected model and Latency Percentiles to understand whether the increase affects typical requests or only the slower requests.

You can then compare models using Model Performance Comparison and use Prompt Analysis to investigate individual interactions. For deeper analysis, the Trace view displays the complete application trace containing the selected LLM event, helping you correlate the LLM request with preceding and subsequent operations and identify potential performance bottlenecks. The Logs tab provides additional information related to the affected requests.

Outcome

By correlating model-level performance metrics with individual LLM interactions, traces, and logs, teams can identify the source of increased latency and troubleshoot LLM-related performance issues more efficiently.

Keep LLM costs in check

When LLM costs increase, use Derived Total Cost and Avg Cost / Request to understand the overall change, identify the contributing models using Cost by Model, and correlate the increase with Input Token Usage by Model, Output Token Usage by Model, and Tokens Per Call by Model.

Get started

Navigate to APM > LLM Observability in your Site24x7 console. Select the required time period and use the Dashboard, Prompt Analysis, and Logs tabs to start analyzing your LLM activity.

For more details, please refer to our LLM Observability help documentation.

Have questions or suggestions? Share them in the comments—we’d love to hear from you!

Happy monitoring,
The Site24x7 team

Like (1) Reply
Replies (0)