Monitoring LLM Systems: Metrics, Logging, Alerting

Michael BrenndoerferFebruary 13, 202646 min read

Part of Language AI Handbook

Monitor production LLM systems with metrics collection, structured logging, dashboard design, alerting rules, and concept drift detection.

Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.

Article links

Make inline references clickable

Monitoring

When a language model leaves the lab and enters production, the work has only just begun. Training a model is a bounded problem: you define a dataset, run experiments, and converge on something that performs well on your evaluation suite. Deploying a model is an unbounded problem: real users send inputs you never anticipated, the world changes beneath your feet, and failures can be silent. A model that returned thoughtful answers in testing might start returning rambling non-answers three months later, and without monitoring you would never know.

Monitoring answers a simple question: is the system doing what you expect, right now? That question turns out to be surprisingly hard to answer for language models. Unlike a relational database query that either returns the right row or throws an error, a language model always returns something. It will confidently answer questions about topics it knows nothing about, shift its tone without warning, or gradually degrade in quality as its context grows stale. These failures are invisible unless you are actively measuring for them.

This chapter covers the full monitoring stack for production LLM systems: the metrics that matter, how to collect and store them, how to build dashboards that expose problems, and how to configure alerts that wake you up when something breaks. Along the way, we will look at the specific observability challenges that language models introduce and work through a practical implementation using OpenTelemetry, Prometheus, and Grafana-style visualizations.

Why LLM Monitoring Differs from Traditional Software Monitoring

Traditional application monitoring focuses on infrastructure health. You watch CPU usage, memory consumption, request latency, and error rates. These signals are objective and binary: a 500-series HTTP response code is an error, a process consuming 100% of CPU is a problem. The signals are also correlated in predictable ways: if your CPU spikes, your latency goes up, and users complain. You can often draw a straight line from symptom to cause.

Language model monitoring requires all of that, plus a layer of semantic monitoring that has no direct analogue in traditional software. A request to an LLM can succeed at the HTTP level while failing completely at the application level. The model returns a 200 OK with a beautifully formatted response that is factually wrong, off-topic, or harmful. You cannot detect this by watching port 80.

The challenges cluster around three core properties of language models:

  • Outputs are probabilistic and hard to measure automatically. A model's output quality is subjective. You cannot write a simple comparison function to check if a response is "correct." You can check for certain patterns (did the response contain a JSON object, did it exceed a length limit), but deep quality assessment requires either human review or a secondary model as a judge.

  • Models degrade silently over time. Input distributions shift as user behavior evolves, new topics emerge, and the context the model was trained on grows stale. This is called concept drift, and it happens gradually. Without longitudinal tracking of quality metrics, you will not notice until the degradation is severe.

  • Failures are multidimensional. A language model can fail in technical ways (timeout, out-of-memory error), semantic ways (irrelevant answer), safety ways (toxic or harmful content), and business ways (response that violates a policy or damages brand reputation). Each dimension requires different instrumentation.

This multidimensionality makes LLM observability a hard engineering problem. In a traditional web service, "is the system healthy?" has a fairly crisp answer: check error rates and latency. For an LLM service, a system can be technically healthy - fast responses, zero HTTP errors, ample GPU headroom - while producing outputs that are quietly misleading users or violating content policies. The monitoring stack has to cover both the mechanical health of the infrastructure and the semantic quality of the outputs it produces.

Consider a concrete scenario: a production customer support bot starts responding with plausible-sounding but incorrect information about a policy that changed six months ago. The bot's latency is fine. The error rate is zero. The infrastructure metrics show nothing unusual. Only if you are tracking automated quality scores or running periodic human review will you catch that the model's answers no longer reflect reality. This kind of silent failure is entirely absent from the playbook for traditional software operations, and it requires a fundamentally different approach to monitoring.

The Monitoring Stack

A complete monitoring system for an LLM service consists of four layers, each building on the previous one.

Instrumentation is the code embedded in your application that generates raw signals: counters, timers, gauges, and log entries that describe what the system is doing at each moment. Good instrumentation is lightweight, consistent, and designed for the specific questions you need to answer. For LLM services, this means tracking the usual HTTP metrics along with token counts, quality scores, and model-specific signals like time to first token.

Collection and storage is the infrastructure that receives those signals and persists them efficiently. Metrics need a time-series database that handles high write throughput and supports range queries (give me the p99 latency for all requests over the last four hours). Logs need a system that can handle large text volume, compress it efficiently, and support full-text and structured queries. Traces need a system that can stitch together spans from multiple services into a coherent call graph.

Visualization is the dashboard layer that renders signals in human-readable form. A good dashboard does not just display numbers; it presents information in a way that makes patterns immediately apparent. Operators who stare at a dashboard for thirty seconds should know whether anything requires attention. This requires careful design choices: which metrics to show, how to normalize them, which time window to default to.

Alerting is the rules layer that automatically notifies operators when signals cross thresholds. Alerting closes the loop between monitoring and action. Without it, monitoring is purely reactive: operators look at dashboards when they suspect a problem. With good alerting, problems surface themselves before users notice them.

These layers are commonly implemented using the open-source LGTM stack (Loki for logs, Grafana for dashboards, Tempo for traces, and Mimir or Prometheus for metrics), though cloud-native alternatives like Datadog, New Relic, or CloudWatch are equally viable. The choice between self-hosted open source and managed cloud services involves tradeoffs in operational burden, cost, data residency, and integration depth. Teams running on Kubernetes often find the open-source stack integrates more naturally with their existing tooling, while teams at early stages often prefer a managed service to avoid the overhead of running yet more infrastructure.

Metrics

Metrics are numeric measurements sampled at regular intervals or computed per event. They are the fastest signal to collect and the easiest to alert on. For LLM services, metrics fall into several natural categories, each measuring a different dimension of system health.

Infrastructure Metrics

Infrastructure metrics are the standard system-level signals that any service should track. They represent the mechanical health of the system independent of the model's output.

  • Request rate: Requests per second arriving at the service. Sudden spikes indicate viral usage or potential abuse. Sudden drops indicate outages. The rate is also the foundation for understanding all other metrics: a latency increase is far more alarming at 500 requests per second than at 5 requests per second.

  • Request latency: Time from request arrival to response completion, measured at p50, p95, and p99 percentiles. The p99 matters more than the mean because LLM latency distributions are heavily right-skewed: a small fraction of requests with very long outputs or high retry counts can have orders-of-magnitude higher latency than the median. If you only monitor the mean, that fraction is invisible. Users hitting those p99 requests experience the worst performance your system can produce.

  • Error rate: Fraction of requests ending in an error. Segment by error type (model timeout, rate limit exceeded, output too long) to distinguish between infrastructure problems and usage patterns. An error rate spike caused by rate limiting from an upstream model provider requires a different response than one caused by your own inference backend running out of memory.

  • Resource utilization: GPU memory and utilization, CPU, RAM, and network I/O. GPU memory deserves special attention because LLM inference is almost always memory-bound; even small memory leaks cause gradual degradation. A model loaded at half the GPU's memory capacity has room for batch growth; one that fills 90% of the GPU has very little margin before out-of-memory errors start occurring.

LLM-Specific Operational Metrics

Beyond infrastructure, LLM services have operational metrics that do not exist in other software systems. These reflect the unique economics and mechanics of token-based language model inference.

Tokens per request is the number of input tokens and output tokens in each request. This is fundamental in a way that has no analogue in traditional services. A SQL query that returns a thousand rows costs roughly the same as one that returns ten rows, in terms of server-side compute. An LLM request with 8,000 input tokens costs far more to serve than one with 100 input tokens, both in time and money. Tracking token counts per request lets you understand your cost structure, identify unusually expensive requests, and set budget guardrails.

Tokens per second (throughput) measures the raw generation speed of the model: output tokens generated per second. This is the benchmark number that most hardware vendors advertise, and it is what ultimately determines the steady-state capacity of your inference backend. Tracking it continuously lets you detect performance regressions after software updates or configuration changes.

Time to first token (TTFT) is the latency from request arrival to the first token appearing in the response. For streaming applications, this is what users feel as "responsiveness," and it is distinct from end-to-end latency. A model that takes two seconds to produce its first token but then streams quickly may feel more responsive than one that takes one second to start but produces tokens slowly. TTFT is dominated by the prefill phase (processing the input prompt), so it scales with input length. Tracking TTFT separately from end-to-end latency lets you distinguish slow prefill from slow generation.

Cache hit rate measures the fraction of requests served from cache. If your system uses KV-cache or semantic caching, cache hits can reduce cost and latency by an order of magnitude. A dropping cache hit rate often signals that query patterns are shifting, which might also predict increases in cost and latency before those increases show up in other metrics.

Context length distribution tracks the distribution of prompt lengths over time. If users are feeding in longer and longer contexts (because they are pasting entire documents, or because conversation history is accumulating), your latency and cost will increase, and you need to know this proactively. Context length also affects memory pressure: longer contexts require more GPU memory for the KV cache during inference, and a sudden shift toward very long contexts can push a system that was comfortably within memory limits into out-of-memory territory.

Quality Metrics

Quality metrics are harder to compute automatically but are the most important for understanding whether your system is doing its job. An LLM service that is fast and reliable but produces bad answers is still a bad service. Quality metrics are what distinguish LLM monitoring from generic API monitoring.

Automated quality scores use secondary models or simpler classifiers to score each response for relevance, coherence, safety, or task-specific properties. For example, a retrieval-augmented generation (RAG) system might score each response for faithfulness to the retrieved documents using a small entailment model. These scores are inherently imperfect but provide continuous coverage that human review cannot match. The key is to calibrate automated scores against human judgments periodically to ensure the scorer remains aligned with the properties you care about.

User feedback signals capture real user judgment at scale. Thumbs up/down ratings, regeneration rates (users who click "try again"), edit rates (users who modify the model's output before using it), and session abandonment rates all proxy for response quality. These signals are noisy - a user might regenerate because of curiosity rather than dissatisfaction - but aggregated across thousands of requests they provide a reliable quality signal that is independent of any model-based evaluator.

Safety classifier outputs track the fraction of responses that trigger or nearly trigger safety filters. If your application has safety requirements, running each response through a classifier and tracking the rate over time lets you detect shifts in user behavior (more adversarial inputs arriving) or model behavior (safety degradation after a fine-tuning update).

The combination of automated scores, user feedback, and safety monitoring provides a three-way view of quality. Automated scores are sensitive but imperfect; user feedback is ground truth but slow and incomplete; safety monitoring catches the tail of unacceptable outputs. Together they give you much higher confidence that the system is performing well than any single signal alone.

Logging

While metrics summarize behavior in numbers, logs capture the raw events. For LLM services, logs are especially valuable because they preserve the actual inputs and outputs, enabling retrospective quality analysis and debugging. When an automated alert fires indicating quality degradation, your first response will be to look at the logs to understand what inputs triggered the problem.

What to Log

A production LLM request log entry should capture enough information to fully reconstruct what happened during the request without needing to contact the user. The essential fields are:

  • A unique request ID (for correlating across logs and traces)
  • Timestamp and latency
  • The model identifier and version
  • Compressed or hashed input (preserving structure while protecting PII)
  • Output text or a hash of the output
  • Token counts (input and output separately)
  • Any metadata relevant to routing (which backend handled the request, whether it was a cache hit)
  • Quality scores from any automated evaluators
  • User feedback if available

Logging the full input and output text is ideal for debugging but raises serious privacy and storage concerns. A practical middle ground is to log full text for a small random sample (say 1-5% of requests) while logging only metadata for the rest. This gives you enough data for quality analysis without storing everything. The sampling can be biased toward unusual requests: low quality scores, high latency, errors, and new user sessions are all good candidates for higher sample rates because they are more likely to contain useful debugging information.

Structured Logging

Logs are most useful when they are structured rather than free-form text. A free-form log line like "Completed request abc123 in 1.2s with 423 tokens" requires regex parsing to extract values. A structured log in JSON format makes the same information immediately queryable:

{
  "request_id": "abc123",
  "latency_ms": 1200,
  "input_tokens": 156,
  "output_tokens": 267,
  "model": "gpt-4o",
  "quality_score": 0.87,
  "timestamp": "2024-01-15T10:23:45Z"
}

This structure enables log aggregation systems like Elasticsearch, Loki, or CloudWatch Logs Insights to run arbitrary queries without custom parsing logic. You can ask questions like "what is the average quality score for requests with more than 2,000 input tokens in the last seven days?" directly in the query interface, without writing a data pipeline first. Structured logging is one of those practices that seems like overhead until you need to debug a production incident under time pressure, at which point it becomes invaluable.

The field naming in structured logs deserves the same care as an API design. Choose names that are descriptive (input_tokens rather than itk), consistent across services (so joins are possible), and stable over time (changing a field name silently breaks all queries that relied on it). Many teams adopt a shared log schema across services to enable cross-service analysis.

Log Retention Strategy

Raw LLM logs can be enormous. A service handling 100,000 requests per day, each with a 1,000-token input and 500-token output, generates gigabytes of text per day. A practical retention strategy tiers logs by freshness and importance:

  • Hot tier (0-7 days): Full logs including text, indexed for fast search. Used for active debugging.
  • Warm tier (7-90 days): Metadata only (timestamps, token counts, quality scores), compressed. Used for trend analysis.
  • Cold tier (90+ days): Aggregated statistics only, archived to cheap storage. Used for auditing and compliance.

The transition between tiers should be automated: your log management system drops text payloads at the warm tier boundary and drops individual records entirely at the cold tier boundary, retaining only rollup aggregations. This keeps storage costs manageable while preserving the data you need for each use case.

Distributed Tracing

Metrics tell you that something is slow. Logs tell you what happened during a specific request. Traces tell you where time was spent across the entire call chain. For a typical LLM application, a single user request might involve several services: a web frontend, an API gateway, a vector database lookup, a prompt assembly service, the model inference backend, and a response post-processing step. Tracing attaches a unique trace ID to the request at entry and propagates it through every service, creating a timeline that shows exactly where latency is coming from.

For LLM services specifically, traces answer questions like: is the latency increase coming from the vector database lookup, the model inference step, or post-processing? Without tracing, you might assume the model is slow when in fact your embedding lookup is the bottleneck. This distinction matters enormously for triage: optimizing the wrong layer wastes engineering time while the real problem continues.

A trace for an LLM request typically consists of several spans. The root span covers the entire request duration from receipt to response. A child span covers the prompt assembly step, showing how long it takes to retrieve context documents and construct the final prompt. Another child span covers the model inference call, showing the time spent in the inference backend. If the model call is to an external API, this span shows the round-trip latency to that provider. Post-processing steps, output validation, and logging each get their own spans.

OpenTelemetry has become the standard instrumentation library for distributed tracing. It provides language-specific SDKs that instrument your code with minimal effort and export trace data to any compatible backend. The key advantage of OpenTelemetry is vendor neutrality: the same instrumentation code can export to Jaeger, Tempo, Zipkin, or any commercial tracing backend, so you are not locked into a specific vendor's SDK. As your monitoring infrastructure evolves, you can change the backend without re-instrumenting your application.

Implementation

Let's build a practical monitoring system for an LLM API service. We will instrument the application with metrics and logging, generate realistic simulated traffic, and demonstrate the kinds of visualizations you would use in a real deployment.

First, install the required packages:

In[3]:
Code
# uv pip install prometheus-client opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp
# uv pip install opentelemetry-instrumentation-fastapi fastapi uvicorn
# uv pip install numpy matplotlib

Core Metrics Instrumentation

We will define a metrics registry using the prometheus_client library. This library exposes metrics in the Prometheus text format, which can be scraped by a Prometheus server or pushed to any compatible backend. For this demonstration, we implement a simplified version of the Prometheus data types so the code runs without a running Prometheus instance.

In[4]:
Code
import numpy as np


# Simulate the prometheus_client API for demonstration
class Counter:
    def __init__(self, name, description, labelnames=()):
        self.name = name
        self._value = 0
        self._labels = {}

    def labels(self, **kwargs):
        key = tuple(sorted(kwargs.items()))
        if key not in self._labels:
            self._labels[key] = {"value": 0}
        return _LabeledCounter(self._labels[key])

    def inc(self, amount=1):
        self._value += amount


class _LabeledCounter:
    def __init__(self, store):
        self._store = store

    def inc(self, amount=1):
        self._store["value"] += amount


class Histogram:
    def __init__(self, name, description, labelnames=(), buckets=None):
        self.name = name
        self._observations = []

    def labels(self, **kwargs):
        return self

    def observe(self, value):
        self._observations.append(value)

    @property
    def mean(self):
        return np.mean(self._observations) if self._observations else 0

    @property
    def p50(self):
        return (
            np.percentile(self._observations, 50) if self._observations else 0
        )

    @property
    def p95(self):
        return (
            np.percentile(self._observations, 95) if self._observations else 0
        )

    @property
    def p99(self):
        return (
            np.percentile(self._observations, 99) if self._observations else 0
        )


class Gauge:
    def __init__(self, name, description, labelnames=()):
        self.name = name
        self._value = 0

    def set(self, value):
        self._value = value

    def inc(self, amount=1):
        self._value += amount

    def dec(self, amount=1):
        self._value -= amount


# Define the metrics registry
REQUEST_COUNTER = Counter(
    "llm_requests_total",
    "Total number of LLM requests",
    labelnames=["model", "status"],
)

REQUEST_LATENCY = Histogram(
    "llm_request_duration_seconds",
    "LLM request latency in seconds",
    labelnames=["model"],
    buckets=[0.1, 0.5, 1.0, 2.0, 5.0, 10.0, 30.0, 60.0],
)

TTFT_HISTOGRAM = Histogram(
    "llm_time_to_first_token_seconds",
    "Time to first token in seconds",
    labelnames=["model"],
    buckets=[0.05, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0],
)

TOKEN_COUNTER = Counter(
    "llm_tokens_total", "Total tokens processed", labelnames=["model", "type"]
)

QUALITY_SCORE = Histogram(
    "llm_response_quality_score",
    "Automated quality score for LLM responses (0-1)",
    labelnames=["model"],
    buckets=[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0],
)

GPU_MEMORY_GAUGE = Gauge("llm_gpu_memory_bytes", "GPU memory currently in use")

Each metric type serves a different purpose. Counters accumulate monotonically and are typically used for totals that only go up (requests, tokens, errors). Histograms record the distribution of a value across configurable buckets, enabling percentile calculations at query time. Gauges represent a point-in-time value that can go up or down, like current GPU memory usage or active connection count.

Simulating Request Traffic

With our metrics defined, we simulate a realistic request workload. The simulation is designed to include a mix of normal traffic and degraded traffic, so we can test whether our alerting correctly detects the quality drop.

In[5]:
Code
import datetime
import random
from dataclasses import dataclass


@dataclass
class RequestMetadata:
    request_id: str
    model: str
    input_tokens: int
    output_tokens: int
    latency_ms: float
    ttft_ms: float
    quality_score: float
    status: str
    timestamp: str


def simulate_request(
    model: str, quality_degraded: bool = False
) -> RequestMetadata:
    """Simulate a single LLM API request with realistic characteristics."""
    request_id = f"req_{random.randint(100000, 999999)}"

    # Input token distribution: most requests are short, some are long
    input_tokens = int(np.random.lognormal(mean=5.5, sigma=1.0))
    input_tokens = min(input_tokens, 8192)

    # Output tokens proportional to input with some noise
    output_tokens = int(input_tokens * random.uniform(0.3, 2.0))
    output_tokens = min(output_tokens, 4096)

    # Latency scales with total token count
    total_tokens = input_tokens + output_tokens
    base_latency = 0.5 + (total_tokens / 1000) * 2.0
    latency_ms = base_latency * 1000 * np.random.lognormal(0, 0.3)

    # Time to first token is faster (prefill phase)
    ttft_ms = latency_ms * random.uniform(0.1, 0.3)

    # Quality score: high normally, degraded if drift is simulated
    if quality_degraded:
        quality_score = random.betavariate(2, 5)  # Right-skewed toward 0
    else:
        quality_score = random.betavariate(8, 2)  # Right-skewed toward 1

    # 2% error rate
    status = "error" if random.random() < 0.02 else "success"

    timestamp = datetime.datetime.utcnow().isoformat() + "Z"

    return RequestMetadata(
        request_id=request_id,
        model=model,
        input_tokens=input_tokens,
        output_tokens=output_tokens,
        latency_ms=latency_ms,
        ttft_ms=ttft_ms,
        quality_score=quality_score,
        status=status,
        timestamp=timestamp,
    )


def record_request_metrics(req: RequestMetadata):
    """Push request metrics to the Prometheus registry."""
    REQUEST_COUNTER.labels(model=req.model, status=req.status).inc()
    REQUEST_LATENCY.labels(model=req.model).observe(req.latency_ms / 1000)
    TTFT_HISTOGRAM.labels(model=req.model).observe(req.ttft_ms / 1000)
    TOKEN_COUNTER.labels(model=req.model, type="input").inc(req.input_tokens)
    TOKEN_COUNTER.labels(model=req.model, type="output").inc(req.output_tokens)

    if req.status == "success":
        QUALITY_SCORE.labels(model=req.model).observe(req.quality_score)


# Generate a week of simulated traffic
np.random.seed(42)
n_requests = 5000
requests_normal = [
    simulate_request("gpt-4o", quality_degraded=False)
    for _ in range(int(n_requests * 0.7))
]
requests_degraded = [
    simulate_request("gpt-4o", quality_degraded=True)
    for _ in range(int(n_requests * 0.3))
]
# Preserve arrival order for the drift dashboard, then shuffle a copy for
# aggregate distributions where temporal order is irrelevant.
timeline_requests = requests_normal + requests_degraded
all_requests = timeline_requests.copy()
random.shuffle(all_requests)

# Record all requests to metrics
for req in all_requests:
    record_request_metrics(req)

# Simulate GPU memory varying over time
gpu_samples = np.random.normal(18.5, 2.0, 1000)  # GB
gpu_samples = np.clip(gpu_samples, 0, 24)

The simulation uses log-normal distributions for token counts and latency because these match the empirical distributions of real production traffic well. Real users mostly send short requests, but a fat tail of long-context requests exists and drives a disproportionate share of total compute. Quality scores are sampled from a beta distribution, which is bounded between 0 and 1, making it a natural fit for scores interpreted as probabilities or fractions.

Out[6]:
Console
Total requests simulated: 5,000
  Successful: 4,905 (98.1%)
  Errors:     95 (1.9%)

Average input tokens:  403
Average output tokens: 450

Quality scores:
  Normal traffic:    0.798
  Degraded traffic:  0.287

Latency percentiles:
  p50:  1550 ms
  p95:  6426 ms
  p99:  12247 ms

The output shows exactly the kind of data a monitoring system would surface. The error rate is within normal range, but the gap between normal and degraded quality scores (roughly 0.8 versus 0.3) is the signal we want our alerting to catch. Notice that the infrastructure metrics (error rate, latency) look completely healthy even while 30% of the traffic is producing poor quality outputs, illustrating why semantic monitoring is essential.

Building Dashboards

Now let's visualize the key metrics as they would appear in a Grafana-style dashboard. The first visualization shows the request latency distribution, which is one of the most informative single charts for understanding system performance.

Out[7]:
Visualization
Histogram of simulated LLM request latency from zero to fifteen thousand milliseconds, with dashed p50, p95, and p99 markers.
Simulated request latency is strongly right-skewed: the median is about 1.6 seconds, while the 95th and 99th percentiles extend much farther into the tail. Percentile markers expose slow requests that a mean alone would hide.

The right-skewed shape is typical of LLM inference. Most requests complete quickly because they have short inputs and short outputs. The long tail comes from requests with thousands of input tokens or that trigger multi-step chain-of-thought generation. When you plot only the mean, these outliers inflate it without revealing the cause of the higher value. Percentile monitoring makes the shape of the distribution visible.

Out[8]:
Visualization
Overlapping histograms of simulated input and output token counts on a logarithmic axis from ten to ten thousand tokens.
Simulated input and output token counts overlap on a logarithmic axis, revealing both the dense body of ordinary requests and the long tail of expensive requests. The tail matters directly for per-request compute cost and memory demand.

The log scale reveals the true spread. On a linear scale, the long tail compresses the majority of the distribution into a thin strip near zero, hiding its structure. On a log scale, you can see that the peak of the input distribution is around a few hundred tokens, but requests extend out to 8,000 tokens. This directly translates to cost variance: a request at the 99th percentile of the input distribution might cost 20-30 times more than one at the median.

Out[9]:
Visualization
Scatter plot of simulated quality scores over time. A shaded region marks degraded traffic after about request 3,430, where the rolling mean drops from about 0.8 toward 0.3 and crosses a 0.6 threshold.
A simulated traffic shift begins after 70 percent of requests. The 100-request rolling mean falls from roughly 0.8 to below the 0.6 alert threshold, making sustained quality degradation visible despite noisy individual scores.

This chart captures what concept drift looks like in practice. The individual scores are noisy, but the rolling mean reveals the trend clearly. An alert tied to the rolling mean crossing the threshold is sensitive enough to catch sustained degradation while ignoring isolated bad responses.

Out[10]:
Visualization
Line chart of GPU memory usage in GB over time with horizontal alert threshold and capacity lines.
Simulated GPU memory usage over 1,000 time samples, centered around 18.5 GB of a 24 GB device capacity. The 22 GB alert threshold provides a roughly 2 GB safety margin before out-of-memory errors begin, giving operators time to investigate before the situation becomes critical.

GPU memory is one of the most important infrastructure metrics for LLM services. Unlike CPU or RAM, GPU memory is a hard limit: when it is exhausted, the inference process crashes rather than slowing down. Setting the alert threshold well below the physical limit (90% rather than 100%) gives operators a warning window to investigate before the system becomes unstable. In practice, this matters especially at traffic peaks, where sudden bursts of long-context requests can push memory usage from comfortable to critical in seconds.

Alerting Rules

Metrics and dashboards help operators understand normal operation and investigate known problems. Alerts are the proactive layer: they tell you when something breaks before users file support tickets.

Effective alerting has two failure modes. Too few alerts means real problems go undetected. Too many alerts means operators get desensitized and start ignoring them, which is equivalent to no alerts at all. The goal is high-signal, low-noise alerts tied directly to user-visible impact.

A practical alert rule specification follows the format used by Prometheus and most modern alerting systems:

In[11]:
Code
# Alert rule definitions (YAML-style config representation)
ALERT_RULES = [
    {
        "name": "HighErrorRate",
        "expr": "rate(llm_requests_total{status='error'}[5m]) / rate(llm_requests_total[5m]) > 0.05",
        "for": "2m",
        "severity": "critical",
        "summary": "Error rate exceeded 5% over the last 5 minutes",
        "runbook": "Check model backend health, review recent deployments",
    },
    {
        "name": "HighP99Latency",
        "expr": "histogram_quantile(0.99, rate(llm_request_duration_seconds_bucket[5m])) > 30",
        "for": "5m",
        "severity": "warning",
        "summary": "p99 latency exceeded 30 seconds",
        "runbook": "Check for requests with very long contexts, review GPU utilization",
    },
    {
        "name": "LowQualityScore",
        "expr": "histogram_quantile(0.5, rate(llm_response_quality_score_bucket[30m])) < 0.6",
        "for": "10m",
        "severity": "warning",
        "summary": "Median quality score dropped below 0.6 over the last 30 minutes",
        "runbook": "Review recent model outputs, check for prompt injection or adversarial inputs",
    },
    {
        "name": "HighGPUMemoryPressure",
        "expr": "llm_gpu_memory_bytes / (24 * 1e9) > 0.90",
        "for": "3m",
        "severity": "warning",
        "summary": "GPU memory utilization exceeded 90%",
        "runbook": "Review context length distribution, consider enabling memory-efficient attention",
    },
    {
        "name": "ModelBackendDown",
        "expr": "up{job='llm-inference'} == 0",
        "for": "1m",
        "severity": "critical",
        "summary": "LLM inference backend is unreachable",
        "runbook": "Check inference server process, review system logs",
    },
]

for rule in ALERT_RULES:
    print(f"Alert: {rule['name']}")
    print(f"  Severity: {rule['severity']}")
    print(f"  Condition: {rule['expr'][:60]}...")
    print(f"  Summary: {rule['summary']}")
    print()
Out[12]:
Console
Alert: HighErrorRate
  Severity: critical
  Condition: rate(llm_requests_total{status='error'}[5m]) / rate(llm_requ...
  Summary: Error rate exceeded 5% over the last 5 minutes

Alert: HighP99Latency
  Severity: warning
  Condition: histogram_quantile(0.99, rate(llm_request_duration_seconds_b...
  Summary: p99 latency exceeded 30 seconds

Alert: LowQualityScore
  Severity: warning
  Condition: histogram_quantile(0.5, rate(llm_response_quality_score_buck...
  Summary: Median quality score dropped below 0.6 over the last 30 minutes

Alert: HighGPUMemoryPressure
  Severity: warning
  Condition: llm_gpu_memory_bytes / (24 * 1e9) > 0.90...
  Summary: GPU memory utilization exceeded 90%

Alert: ModelBackendDown
  Severity: critical
  Condition: up{job='llm-inference'} == 0...
  Summary: LLM inference backend is unreachable

Each alert rule includes a for clause that requires the condition to be true for a sustained period before firing. This prevents transient spikes from generating noise. The runbook field links each alert to documented response procedures, transforming an alarm from a puzzle into an action item. An alert that fires at 3 AM is far more useful if it comes with a clear description of what to check and what to do, rather than just a metric value and a name.

The for duration deserves careful tuning. For a hard outage (model backend unreachable), even a one-minute window is too slow: users are already experiencing errors. For a quality degradation signal, a ten-minute window is appropriate because quality scores fluctuate and a sustained drop is much more meaningful than a momentary dip. Match the for duration to the expected dynamics of the underlying condition.

Simulating Alert Detection

Let's verify that our quality metric alert would correctly fire on the degraded traffic scenario we simulated:

In[13]:
Code
def evaluate_quality_alert(requests_window, threshold=0.6):
    """Evaluate whether the quality alert condition is met."""
    successful = [r for r in requests_window if r.status == "success"]
    if not successful:
        return False, None

    scores = [r.quality_score for r in successful]
    median_score = float(np.median(scores))

    return median_score < threshold, median_score


# Evaluate alert on normal traffic window
alert_fired_normal, score_normal = evaluate_quality_alert(requests_normal[:500])

# Evaluate alert on degraded traffic window
alert_fired_degraded, score_degraded = evaluate_quality_alert(
    requests_degraded[:500]
)
Out[14]:
Console
Quality Alert Evaluation:

Normal traffic window:
  Median quality score: 0.815
  Alert fired: False

Degraded traffic window:
  Median quality score: 0.263
  Alert fired: True

Status: ALERT FIRING

The alert correctly remains silent during normal operation and fires during the degraded period. This demonstrates that the threshold is calibrated to capture quality issues without generating false positives on healthy traffic. In a real system, you would tune this threshold by analyzing historical quality score distributions and picking a value below the natural variance of normal operation but above the range seen during known degradation incidents.

Structured Log Emission

The final piece is logging. Here we implement a simple structured logger that would integrate with a log aggregation system like Loki or Elasticsearch:

In[15]:
Code
class LLMRequestLogger:
    """Emits structured JSON logs for LLM requests."""

    def __init__(self, service_name: str, sample_rate: float = 0.05):
        self.service_name = service_name
        self.sample_rate = sample_rate
        self._log_buffer = []

    def log_request(self, req: RequestMetadata, include_text: bool = False):
        """Log a completed request."""
        # Always log metadata; only log text for sampled fraction
        should_log_text = random.random() < self.sample_rate

        entry = {
            "service": self.service_name,
            "request_id": req.request_id,
            "timestamp": req.timestamp,
            "model": req.model,
            "status": req.status,
            "latency_ms": round(req.latency_ms, 2),
            "ttft_ms": round(req.ttft_ms, 2),
            "input_tokens": req.input_tokens,
            "output_tokens": req.output_tokens,
            "quality_score": round(req.quality_score, 4),
            "text_sampled": should_log_text,
        }

        self._log_buffer.append(entry)
        return entry


logger = LLMRequestLogger(service_name="language-api", sample_rate=0.05)

# Log a sample of requests
for req in all_requests[:20]:
    logger.log_request(req)

# Show a few representative log entries
text_sampled_count = sum(1 for e in logger._log_buffer if e["text_sampled"])
Out[16]:
Console
Sample log entries:
------------------------------------------------------------
{
  "service": "language-api",
  "request_id": "req_897003",
  "timestamp": "2026-08-15T02:44:39.339861Z",
  "model": "gpt-4o",
  "status": "success",
  "latency_ms": 1014.35,
  "ttft_ms": 274.43,
  "input_tokens": 35,
  "output_tokens": 39,
  "quality_score": 0.6873,
  "text_sampled": false
}

{
  "service": "language-api",
  "request_id": "req_376185",
  "timestamp": "2026-08-15T02:44:39.325797Z",
  "model": "gpt-4o",
  "status": "success",
  "latency_ms": 9846.49,
  "ttft_ms": 2896.58,
  "input_tokens": 1143,
  "output_tokens": 1304,
  "quality_score": 0.9424,
  "text_sampled": false
}

{
  "service": "language-api",
  "request_id": "req_488069",
  "timestamp": "2026-08-15T02:44:39.340554Z",
  "model": "gpt-4o",
  "status": "success",
  "latency_ms": 1703.85,
  "ttft_ms": 216.62,
  "input_tokens": 278,
  "output_tokens": 176,
  "quality_score": 0.2353,
  "text_sampled": false
}

Log buffer: 20 entries
Text-sampled entries (5% rate): 1

Each log entry is a self-contained JSON object with all fields needed for debugging, cost analysis, and quality auditing. The text_sampled flag indicates whether the full prompt and response text were captured, enabling targeted retrieval of sampled records during investigations. Notice that even without the full text, a single log entry tells you a great deal: which model handled the request, how long it took, how many tokens it consumed, and what the automated quality score was.

Dashboards

A monitoring dashboard should answer three questions at a glance: Is the system up? Is it performing well? Is the output quality acceptable? Everything else is detail that should be available on drill-down.

A practical LLM service dashboard organizes panels into three rows. The top row shows the health signals that require immediate attention: request rate, error rate, and p99 latency. If any of these is anomalous, an operator knows immediately that something requires investigation. The second row shows resource utilization: GPU memory and compute, queue depths, and active connection counts. These inform capacity planning and root cause analysis. The third row shows quality signals: rolling quality scores, safety classifier hit rates, and user feedback metrics. These are slower-moving signals but often the first to reveal model drift.

Dashboard design benefits from two principles borrowed from industrial process control. The first is "management by exception": normal is not interesting. Configure charts to show the deviation from baseline rather than the absolute value. A latency chart that shows "300 ms above normal" is more actionable than one showing "1,500 ms absolute," because the reader does not need to remember what "normal" looks like to understand that something is wrong. The second principle is information hierarchy: aggregate at the top, disaggregate on demand. Start with a single number for each metric and let operators click through to the distribution or the individual request trace. Dashboard clutter is a form of alert fatigue: if operators have to scan 40 panels to assess system health, they will either miss problems or stop looking at the dashboard altogether.

The time range selector is one of the most important controls on a dashboard. Short windows (last five minutes) reveal transient spikes and active incidents. Medium windows (last 24 hours) show intraday patterns like traffic peaks during business hours. Long windows (last 30 days) reveal trends: gradual latency increases, slowly declining quality scores, or growing token usage that predict future capacity problems. A good dashboard defaults to the last one or two hours for operational monitoring but makes it easy to zoom out for trend analysis.

Grafana's built-in alert annotations are one of its most useful features for LLM monitoring. When an alert fires, Grafana can mark the corresponding time on every panel in the dashboard with a colored annotation line. This makes it trivial to correlate an alert with the underlying metrics: you can see immediately whether the error rate spike preceded the quality score drop or followed it, which often points toward the root cause.

Alerting Philosophy and Practices

The most important property of a good alert is that it requires action. An alert that operators routinely acknowledge and dismiss without doing anything is a broken alert. It trains operators to ignore the system and erodes trust in the entire monitoring stack. Before adding an alert, ask: if this fires at 3 AM, what action would the on-call engineer take? If the honest answer is "look at it, see nothing actionable, and go back to sleep," the alert should either be converted to a non-paging notification or removed entirely.

This principle has a name in the SRE literature: alerts should be symptoms, not causes. An alert on high CPU usage is a cause-based alert: it tells you something about the internals of the system but not whether users are affected. An alert on high p99 latency is a symptom-based alert: it directly measures what users experience. Prefer symptom-based alerts for your paging tier and use cause-based alerts only for warning-level notifications that inform investigation rather than demand immediate response.

For LLM services, a practical tier of alerts looks like the following. Critical alerts, which wake someone up, cover complete service outages (model backend unreachable), catastrophic error rates above 10%, and safety classifier failures where the monitoring infrastructure itself is down. Warning alerts, which notify but do not wake, cover error rates between 2-10%, p99 latency above service-level objectives, and quality score degradation below the defined threshold. Informational alerts cover cost anomalies (token consumption significantly above baseline) and slow quality drift that does not yet require action but should be tracked.

Alert thresholds deserve the same rigor as any other engineering decision. A threshold that is too tight generates noise; one that is too loose misses real problems. The right starting point is usually to set thresholds based on historical distributions: alert when a metric falls more than two or three standard deviations from its 30-day rolling mean. This adapts to seasonal patterns and gradual growth without requiring manual tuning. Absolute thresholds (error rate above 5%) are simpler to understand but require re-tuning as the system evolves.

Runbooks are the underrated half of good alerting. An alert fires. An on-call engineer wakes up. What do they do? A well-written runbook describes the alert in plain English, explains what the metric means, lists the most common causes (in order of frequency), describes the diagnostic steps to take first, and provides links to relevant dashboards and commands. A runbook written when you set up the alert is far more useful than one written from memory at 2 AM during an incident. Teams that invest in runbooks find that their mean time to resolution drops significantly, because junior engineers can handle incidents that would previously have required escalation to senior staff.

Concept Drift and Long-Term Quality Monitoring

Concept drift is the gradual divergence between the world as the model knows it and the world as it currently is. A model trained on data through 2024 will gradually give less accurate answers about events and developments that emerged afterward. More subtly, language itself drifts: slang evolves, domain-specific terminology shifts, user expectations change. A model that was excellent at answering questions in 2024 might be only adequate by 2026, not because the model changed but because the world it describes changed.

For most applications, concept drift is slow enough that it will not trigger any operational alert for months. Error rates stay flat. Latency stays flat. Users may notice a gradual decline in quality, but the gradual nature means few will file support tickets. The degradation accumulates silently until it is severe, at which point fixing it requires significant effort: either a full retraining run or a substantial retrieval augmentation project.

Detecting concept drift requires longitudinal quality tracking, not just point-in-time monitoring. The patterns to watch for are:

  • Gradual quality decline: The rolling quality score decreases by 5-10% over weeks or months. No single day looks alarming, but the trend is clear over longer windows. A weekly quality report that computes the mean score for the past seven days and compares it to the same metric 30 and 90 days ago is a simple but effective early warning system.

  • Input distribution shift: The distribution of prompt topics, lengths, or vocabulary drifts away from the training distribution. This can be detected by embedding incoming prompts with a frozen encoder model and tracking the centroid of the embedding distribution over time. If the centroid drifts more than a set distance from the original training distribution centroid, that is evidence that users are asking about topics the model was not trained on.

  • Error mode shift: The types of errors the model makes change over time. Early in deployment, most errors might be minor factual inaccuracies. Over time, if the world has changed, the model might make more systematic errors on topics where the ground truth has shifted. Tracking error categories (factual error, incoherence, irrelevance, safety violation) separately provides much richer signal than a single quality score.

Addressing concept drift requires either retraining on newer data or retrieval-augmented generation, where the model can access current information through a search index. Monitoring gives you the evidence you need to make that investment at the right time. Without monitoring, teams typically discover the need for retraining when users complain loudly enough, which often means the degradation has been ongoing for months. With monitoring, you can schedule retraining proactively based on objective quality metrics, which produces better outcomes and less organizational friction.

There is an important distinction between concept drift and distribution shift. Distribution shift refers to changes in the statistical distribution of inputs: users are now asking longer questions, or sending inputs in a different language, or covering a topic that was rare before. Concept drift refers specifically to the model's knowledge becoming stale. Both cause quality degradation, but they require different responses. Distribution shift might be addressed by adding few-shot examples, adjusting the system prompt, or fine-tuning on new user data. Concept drift requires updating the model's knowledge, either through retraining or retrieval augmentation. Tracking both input distributions and output quality gives you the information to distinguish between them.

Cost Monitoring

Cost is a dimension of production health that is easy to overlook until a surprise bill arrives. Language models charge by the token, and token consumption can vary by orders of magnitude depending on user behavior. A simple cost monitoring strategy tracks three things: total token expenditure per day, the distribution of tokens per request, and the fraction of total cost contributed by each user segment or use case.

Total daily expenditure is the headline metric. Set a budget alert at 150% of your expected daily spend. If costs spike above that threshold, investigate immediately: it might indicate a loop in your application code that is calling the model in a tight loop, an abuse pattern where a user is submitting very large inputs, or legitimate usage growth that requires a capacity planning conversation.

The distribution of tokens per request reveals whether high costs are driven by many average-cost requests or by a few very expensive outliers. If a small fraction of requests (say the top 1%) accounts for 30-40% of total cost, you may want to add token limits or routing rules to protect against the most expensive inputs. This is a business decision as much as a technical one: very long-context requests might be your highest-value users, or they might be abusers circumventing rate limits.

Cost attribution by use case or user segment is important in multi-tenant or multi-feature applications. If your product has a document summarization feature and a chat feature, attributing costs separately lets you understand which feature is economically viable and which is a cost center. This is also essential for accurate product pricing: if you are charging users a flat fee, you need to know whether your cost per user is consistent across the user population or whether a small segment of heavy users is driving costs that cannot be covered by the flat fee.

Sampling Strategies for Production Traffic

At scale, it is neither practical nor necessary to analyze every request with the same depth. A tiered sampling strategy concentrates coverage where it matters without the storage and compute overhead of full analysis everywhere.

Uniform random sampling at a low rate (1-5%) provides a statistically representative sample for quality analysis and trend monitoring. If you sample 2% of requests uniformly, you can estimate the population mean quality score with high confidence given sufficient volume. For a service handling 100,000 requests per day, 2% sampling yields 2,000 samples, which is more than enough for reliable statistical analysis.

Stratified sampling over-represents the edges of the distribution. Sample at a higher rate for requests with very long contexts (where quality is more variable), very short contexts (where quality might degrade differently), and requests from new users (who are most sensitive to first impressions). Stratified sampling lets you allocate your analysis budget where it provides the most signal.

Anomaly-triggered sampling captures requests that look unusual according to some fast heuristic: quality score below a threshold, latency in the top 1%, or input length in the top 0.1%. These are the requests most likely to contain interesting failure modes. Saving all anomalous requests rather than just a fraction ensures that rare but important failure modes are not lost.

Feedback-triggered sampling captures all requests that received explicit negative feedback from users. If a user clicks "bad response" or "regenerate," that request should be saved in full for analysis. These are your ground-truth labels for quality, and every one is valuable.

Limitations and Practical Challenges

The most significant limitation of LLM monitoring is the fundamental difficulty of automatically measuring output quality. Almost every metric discussed in this chapter is either a proxy for quality (latency, error rate) or requires a secondary model to compute (automated quality scores). Secondary models bring their own accuracy limitations: a quality scoring model that gives high scores to plausible-sounding but wrong answers provides false assurance. The only ground truth for output quality is human judgment, which does not scale to production traffic volumes.

Practical teams address this through sampling strategies combined with calibration loops. Rather than trying to evaluate every request, they maintain a human review queue that samples a small fraction of production traffic, with intelligent sampling that over-represents edge cases, low-confidence predictions, and new input types. Human-reviewed samples serve two purposes: they provide direct quality signal for the sampled requests, and they feed back into training data for the automated quality scorer. Over time, the scorer becomes better calibrated to the specific quality dimensions that matter for that application.

Privacy is another real constraint. Logging full request and response text exposes potentially sensitive user data. Organizations subject to GDPR, HIPAA, or similar regulations need to think carefully about what they can log, how long they can retain it, and who can access it. In practice, many teams log only metadata and hashes for the vast majority of requests, with full text logging restricted to consented opt-in test users or synthetic requests. This restriction limits the depth of quality analysis possible from logs, which is a real operational cost of privacy compliance. Some teams handle this by having users explicitly opt into full logging in exchange for a quality guarantee or other benefit.

Monitoring infrastructure itself can fail. If your Prometheus scraper falls behind, you lose metric data. If your Loki ingest pipeline is overloaded, log entries are dropped. If your alerting system has a bug, alerts do not fire. Production monitoring requires its own monitoring: track the health of the collection and storage layer, alert on ingest lag, and verify alerts by periodically injecting known anomalies into the system (a practice sometimes called "chaos testing for observability"). An unmonitored monitoring system is worse than no monitoring system, because it gives you false confidence.

Finally, monitoring adds overhead. Each instrumentation call, each metric update, and each log write consumes CPU cycles and I/O. For high-throughput services, the monitoring infrastructure can become a bottleneck. The standard mitigation is to do most monitoring work asynchronously: record metrics to a local in-process buffer and flush to the collector in a background thread, rather than blocking the request path for each metric update. Log writes should go to a local buffer that is flushed periodically or when it fills, not synchronously on each request. Tracing spans should be sampled at the trace level (keep 1% of traces end-to-end) rather than the span level, to avoid the overhead of recording every span.

Summary

Monitoring turns LLM deployment from a one-time event into an ongoing feedback loop between the system and its operators. The key concepts covered in this chapter are:

  • LLM monitoring requires semantic as well as operational metrics. Traditional infrastructure metrics (latency, error rate, resource utilization) are necessary but not sufficient. Quality metrics, safety monitoring, and input distribution tracking are unique to language model services and cannot be omitted without leaving significant blind spots.

  • The monitoring stack has four layers: instrumentation in application code, collection and storage in time-series databases and log aggregators, visualization in dashboards, and alerting through rules that notify operators when thresholds are exceeded. Each layer depends on the ones beneath it, and weaknesses in any layer propagate upward.

  • Tokens are the fundamental unit of LLM economics. Tracking tokens per request, tokens per second throughput, and time to first token separately gives you a full picture of both cost structure and user-facing performance that latency alone cannot provide.

  • Structured logging provides the raw material for debugging and retrospective quality analysis. JSON-formatted logs with consistent field names enable powerful ad-hoc queries and long-term trend analysis without requiring a custom data pipeline for each new question.

  • Effective alerting requires high signal-to-noise ratio. Alerts should be tied to user-visible impact and always imply a specific action. Alert fatigue is as dangerous as insufficient alerting: an on-call team that ignores alerts is effectively unmonitored. Runbooks transform alerts from alarms into documented procedures.

  • Concept drift is a slow, silent failure mode. Longitudinal quality tracking over weeks and months is essential for catching gradual degradation before it becomes severe. Distinguishing concept drift from distribution shift helps direct the appropriate remediation.

  • Automated quality scoring is imperfect but scalable. Combining automated scoring for coverage with human review for calibration gives the best of both approaches: the breadth of full-traffic coverage and the accuracy of human judgment.

  • Cost monitoring deserves equal status with quality and performance monitoring. Token-based pricing means that usage patterns directly drive costs, and an unmonitored application can generate surprising bills from a small number of outlier requests.

The monitoring patterns in this chapter extend naturally to the other operational concerns covered in this part of the book. Auto-scaling decisions in the next chapter depend on the same request rate and latency metrics collected here. Model routing strategies use the quality scores and error rates tracked by this monitoring system. Without monitoring, the entire production system operates blind, and the investments made in model quality, inference optimization, and system architecture cannot be properly evaluated or maintained over time.

Quiz

Ready to test your understanding? Take this quick quiz to reinforce what you've learned about monitoring LLM systems.

LLM Monitoring Quiz

Question 1 of 80 of 8 completed
Which metric is most useful for measuring user-perceived responsiveness in a streaming LLM application?

Comments

No comments yet. Be the first to share your thoughts!

Reference

Citation details

Cite or share this article.

BIBTEXAcademic
@misc{brenndoerfer2026monitoringllm, author = {Michael Brenndoerfer}, title = {Monitoring LLM Systems: Metrics, Logging, Alerting}, year = {2026}, url = {https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production}, organization = {mbrenndoerfer.com}, note = {Accessed: 2026-09-30} }
APAAcademic
Michael Brenndoerfer (2026). Monitoring LLM Systems: Metrics, Logging, Alerting. Retrieved from https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production
MLAAcademic
Michael Brenndoerfer. "Monitoring LLM Systems: Metrics, Logging, Alerting." 2026. Web. September 30, 2026. <https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production>.
CHICAGOAcademic
Michael Brenndoerfer. "Monitoring LLM Systems: Metrics, Logging, Alerting." Accessed September 30, 2026. https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production.
HARVARDAcademic
Michael Brenndoerfer (2026) 'Monitoring LLM Systems: Metrics, Logging, Alerting'. Available at: https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production (Accessed: September 30, 2026).
SimpleBasic
Michael Brenndoerfer (2026). Monitoring LLM Systems: Metrics, Logging, Alerting. https://mbrenndoerfer.com/writing/monitoring-metrics-alerting-logging-dashboards-llm-production

About the author

Continue with the full handbook

This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.

Explore Language AI Handbook
Newsletter

Stay up to date

Get articles, book updates, and news delivered to your inbox.

No spam, unsubscribe anytime.

or

Join the community

Sign in to remove popups, track your reading progress, and join the discussion.