Architecture17 min

Telemetry without the mystery: Prometheus, Loki, Tempo and Grafana, and why OpenTelemetry

By Dorian Chávez · founder of Hábil and integration architect ·

What Prometheus, Loki, Tempo and Grafana answer, how metrics, traces and logs connect, and when to read logs with an agent or instrument with OpenTelemetry.

It is Monday, nine in the morning, and a bank's transfers are taking three times longer than usual. Complaints come in, an incident is opened, and five teams sit down on the same call: digital channels, the core system, antifraud, networks and infrastructure. Each one checks its own screen and each one reaches the same conclusion: "on our side everything is fine." Two hours go by before someone notices that the slowness was in a query to a single service.

That scenario, illustrative and with no customer behind it, is not about a bad team. It is about a system in which nobody can see the complete path of an operation. Telemetry is the answer to that problem: the data a system emits about itself while it works, so that it can be understood from the outside, without opening it up and without guessing. This article is for whoever is responsible for those systems in a bank, a fintech, an insurer, a retail chain or a regulated company, and also for the business director who wants to understand what is being asked of them when the technical team says "we need observability."

We will talk about four widely used open tools —Prometheus, Loki, Tempo and Grafana— and one standard, OpenTelemetry, that ties them together. We will explain what each one answers, how they cross-reference each other, why a common standard makes sense, how they are set up in containers and in Kubernetes, when reading the logs is enough and when the software has to be instrumented from the inside, and what all of this costs and what risks it brings.

Three signals for three different questions

A system in production can tell its story in three ways, and each one answers a different question. This article focuses on the three most common operational signals, which OpenTelemetry calls signals [2]: metrics, traces and logs.

  • Metrics: what is happening, and how much? A metric is a measurement captured while the system runs: how many transfers per minute, how long 95% of them took, how many failed. It is a number over time, cheap to store and excellent for spotting a trend or triggering an alert. It doesn't say which operation failed.
  • Traces: where did this operation go and where did it stop? A trace is the complete path of a request through all the services it touched, with the time it spent in each [2]. It is the map of a single operation. Each segment of that map is called a span.
  • Logs: what did the system say at that moment? A log is the record of an event: a line of text or data with the date, the severity and the message. It is the fine detail: "the connection pool timed out."

None of the three is enough on its own. The metric warns but doesn't locate; the trace locates but doesn't explain; the log explains but, without context, is a needle in a haystack. The good practice, and the thread of this article, is to use them as a chain. In a well-designed instrumentation, the metric usually warns, the trace usually narrows down the path and the logs can supply the detail; the quality of the answer depends on what was emitted and retained.

What each tool answers

Each signal has an open tool designed to store and query it. Before the table, a clarification that helps a committee: none of these tools is interchangeable with another. They are specialized stores.

What each tool answers
ToolWhat it storesQuestion it answersHow it is queriedWhat it is not
PrometheusMetrics: numbers over timeIs latency going up? How many errors per minute?PromQL, e.g. rate(http_requests_total[5m]) [9]A system for billing or reconciling with accounting-grade accuracy [9]; nor a log store
LokiLogsWhat did service X say when it failed?LogQL: you choose the stream and filter the text [15]A full-text search engine: it doesn't index the content of the line, only labels [15]
TempoTracesIn which service did the time of this operation go?TraceQL, with a syntax similar to the other two [18]An automatic generator of business metrics
GrafanaNothing: it queries the othersWhat do I see, and how do I jump from one piece of data to the other?Dashboards, exploration and alerts over any source [19]The system that captures or keeps the telemetry

Two clarifications that save misunderstandings. The first: Prometheus is not a cash register. Its own documentation warns that, if total accuracy is needed, such as per-request billing, it is not a good option [9]. It is for operating —knowing whether the system is healthy—, not for reconciling. In a bank or a fintech that boundary matters: reconciliation lives in the accounting records, not in a latency dashboard.

The second: Loki is cheap precisely because it doesn't index the content of the logs, only a few labels per stream (which service and which environment it comes from). The data is compressed into chunks and stored in object storage [15]. The trade-off is that searching for text inside the logs is done at query time, not with a prior index, and that labels have to be chosen carefully. We will come back to that in the costs.

The role of Grafana

Grafana is the window. According to its documentation, it lets you query, visualize, alert on and explore metrics, logs and traces regardless of where they are stored [19]. Each store is connected as a data source: Prometheus, Loki and Tempo are three sources, and Grafana can have others alongside, such as SQL databases [20].

Its value is not in the pretty dashboard, but in that it cross-references the sources. From one place you can go from a chart to a trace and from a trace to its logs. For an executive, that translates into something concrete: Grafana can bring the investigation into one view if sources, permissions and links are configured, instead of five screens and a call.

How the information is cross-referenced

That the three signals exist doesn't mean they are connected. The connection doesn't appear by installing Grafana: it is designed from the moment the software is instrumented and tested end to end. The common thread is an identifier, the trace_id, which travels with each operation. The propagation standard that OpenTelemetry uses by default, W3C Trace Context, carries that identifier in a header called traceparent from one service to the next [2].

Imagine the trace_id as the tracking number of a package. If each service stamps it on what it produces, afterwards you can ask for "everything that happened with tracking number 4F2A…".

From metric to trace: exemplars

An exemplar is a specific sample attached to an aggregated metric. Think of a latency chart: each point summarizes thousands of operations. An exemplar is, at that point, the example of a real operation —with its trace_id— that contributed to that measurement. In Grafana it appears as a star on the chart; when you click it, the trace of that operation opens in Tempo [20].

For the click from the chart to the trace to work, four pieces must be switched on; if one is missing, the click leads nowhere. Grafana's documentation sums it up like this: metrics give the aggregated view and traces give the fine-grained view of a request [20]. It is worth knowing the conditions before promising it to anyone:

  1. The application has to emit the trace_id together with the measurement, inside a trace that is being kept.
  2. Prometheus has to store the exemplars. Today that capability is enabled with a flag marked as experimental and is disabled by default [12].
  3. In Grafana you have to configure the Prometheus source so that the link points to Tempo and says which field carries the trace_id [20].
  4. The referenced trace still has to exist: if sampling or retention has already discarded it, the click leads nowhere.

An alternate route is for Tempo to compute metrics from the traces —request rate, errors and duration, the well-known RED trio (Rate, Errors, Duration), the three standard measurements used to watch a service— and send them to a Prometheus-compatible database, with their exemplars [17]. That route must be tested against the limits on active series, a topic that is taken up again in the costs.

From trace to log: the `trace_id` on every line

From a trace in Tempo, Grafana can open the logs of that service in that time window, querying Loki. That is configured in the Tempo data source (trace to logs), and an equivalent link to metrics exists (trace to metrics) [21]. To correlate a log with its trace, include the trace_id; add the span_id when you need to narrow the search to the span that emitted the log. OpenTelemetry's logs data model provides for them as dedicated fields [40], and for formats that are not OTLP it recommends naming them trace_id, span_id and trace_flags [41].

From log to trace

The reverse path is configured in Loki with derived fields: a rule that extracts the trace_id from the line and turns it into a link to Tempo [47]. Both sides are necessary: configuring only Tempo doesn't make a log navigable to its trace.

A golden rule of this cross-referencing: the trace_id, and in general any customer, order, account or policy identifier, must not be a Loki label. It should travel inside the content of the log, or as structured metadata. Further on we explain why [14].

Why OpenTelemetry and not each technology directly

The four tools above have their own ways of receiving data. Each application could be programmed to talk directly to each one: one library for Prometheus metrics, another for logs to Loki and another for traces to Tempo. It works, but it ties each service's code to three different models, three retry strategies and three migration paths.

OpenTelemetry (abbreviated OTel) is a CNCF project, the foundation that hosts Kubernetes and Prometheus. It defines itself as a framework for generating, exporting and collecting telemetry; it is vendor-neutral and, importantly, it is not a backend: it doesn't store or draw anything [1]. It provides four pieces:

  • APIs and SDK: per-language libraries with which the software emits metrics, traces and logs with a common model.
  • OTLP: the common protocol for transporting the three signals, over gRPC (port 4317 by default) or HTTP (4318) [4].
  • Semantic conventions: standard names for what is measured, so that "the service" or "the pod" are called the same in all three signals [46]. Without that agreement, the cross-referencing between tools breaks.
  • Collector: an intermediate program that receives the telemetry, enriches it, filters it and forwards it to one or several destinations [6].

The practical consequence: the application emits once, in a standard language, and the decisions about where it goes, what is filtered and what is hidden live in a controlled configuration, not scattered through each team's code. The stores themselves already speak that language: Loki, Tempo and Alloy can receive OTLP, and Prometheus can receive OTLP metrics when --web.enable-otlp-receiver is explicitly enabled [13][16][18][26].

For an executive, the argument is one of portability and control. OpenTelemetry's documentation frames it as not being tied to a vendor and learning a single set of conventions [1]. OpenTelemetry reduces the code's dependence on the destination, although a migration may still require validating conventions, exporters, sampling, dashboards, alerts and retention; the destination is changed in the Collector.

Two caveats in the interest of honesty, because this kind of promise gets inflated easily:

  • Portability is not free. Even with OpenTelemetry, the dashboards, alerts and queries written in PromQL, LogQL and TraceQL are rewritten if the store changes. What is saved is re-instrumenting the software, which is usually the most expensive part.
  • Maturity varies by signal and by language. According to the specification, traces, metrics and logs have a stable protocol, but the metrics SDK is listed as "mixed" and profiles are still in development [3]. By language, Java has all three signals stable; Go, Python and JavaScript have stable traces and metrics, while logs are at release candidate in Go and in development in Python and JavaScript [5]. Whoever programs in Java starts from firmer ground; whoever programs in other languages should check the state of logs before betting on them.

As a reference for adoption, in the CNCF's 2024 annual survey, in the question on incubated projects (n=689), 39% reported OpenTelemetry in production and 23% in evaluation; it is a sample of its community, not market share [44].

How it is set up: containers and Kubernetes

Telemetry needs three things along the way: something to generate it (the software), something to collect and process it (the Collector or an agent) and something to store and display it (Prometheus, Loki, Tempo, Grafana). Where the collector is placed changes depending on the platform.

In Docker and Docker Compose

Compose declares several services, networks and volumes in a single file, and that is why it is used for development, integration testing and small environments [28]. The typical setup is: the applications send OTLP to a container with the Collector (or Alloy); the Collector distributes the traces to Tempo, the logs to Loki and the metrics to Prometheus; and Grafana queries all three. The Collector runs as an official image with its configuration file mounted; without that file it doesn't start [28].

There are two details about logs in Docker that tend to be surprising:

  • Docker captures standard output and standard error by default. If the application writes only to files, docker logs will not show those lines. That is why the official web server images redirect their files to standard output [28].
  • The default driver doesn't rotate logs. The json-file driver has no size limit by default, and Docker recommends the local driver, which rotates by default (according to Docker, 20 MB files, up to five) [28]. A disk full because of logs is a real operations incident.

In addition, the delivery mode matters: in blocking mode, which is the default, a slow driver can slow down the application; in non-blocking mode it uses a buffer and, if it fills up, drops messages [28]. And with the Fluentd driver in synchronous mode, if there is no connection the container stops [29]. You have to choose deliberately between not losing logs and not slowing down the service.

In Kubernetes

Kubernetes recommends that applications write to standard output and that log storage be external, with a lifecycle separate from the pod; it doesn't ship log storage of its own [27]. Rotation is done by the kubelet (Kubernetes default: 10 MiB files and five per container) and kubectl logs only sees the most recent file [27]. That is, without a collector, the logs of a pod that restarted or rotated can be lost.

Kubernetes' documentation describes three patterns for collecting at the cluster level: an agent per node, an agent in each pod, or the application sending directly to the destination [27]. OpenTelemetry translates them into these setups:

How it is set up: containers and Kubernetes
PatternWhat it isWhen it fitsWhat it costs
DaemonSet (agent per node)A collector on each machine of the cluster, which reads the logs of that node's containers and receives OTLP from nearby applicationsStandard-output logs, node metrics; it is the preferred pattern for reading node logs [30][31]Host permissions and mounts; if two collectors read the same files, they duplicate data [31]
SidecarA collector as an additional container inside the podStrong per-application isolation or a special local needMultiplies consumption and configuration in every pod [7]
Gateway (central Deployment)Central collectors that receive from the agents and apply common policies: filtering, sampling, credentialsCentralized controls and output to several destinations [8]"One more thing to maintain that can fail"; adds latency and cost [8]
OperatorA controller that manages the collectors and can inject automatic instrumentation into the podsStandardizing instrumentation across many servicesRequires cert-manager; changes the pod and requires a restart [34]
Grafana AlloyGrafana's distribution of the OpenTelemetry Collector, with native Prometheus and Loki support [23]A single agent for metrics, logs and traces toward the Grafana stack [26]; for pod logs it uses the Kubernetes API or reads node files, and the DaemonSet is the required mode for pod logs [24][25]It belongs to a vendor; it doesn't replace data governance [23]. It is deployed as a DaemonSet, StatefulSet or Deployment depending on the task [25]

Four precautions that the documentation points out and that, when ignored, are paid for in rework during implementation:

  1. Label each log with its owner. A log read from a file doesn't know which pod it comes from. The k8sattributes processor attaches the pod, the namespace, the deployment and the node to each record, and it is stable for all three signals [33]. File reading (filelog) is in beta for logs [32] and, by default, only reads what arrives after it starts; without saving its position, a restart can duplicate or lose lines [32].
  2. One metric, a single writer. Two collectors reporting the same series produce out-of-order data in Prometheus [8].
  3. The collector fails too. With a sending queue, retries and persistent storage configured, it can withstand transient failures; even so, a full disk, a prolonged outage or retry limits can cause loss [35]. It is sized and watched like any other service.
  4. With the Operator, order matters. Automatic instrumentation requires its configuration resource to exist before the pod, the annotation in the right place and a restart; it also overwrites variables such as JAVA_TOOL_OPTIONS [34].

A warning about Promtail. It is the log agent that many teams installed years ago for Loki. Grafana declared its end of life on March 2, 2026: it no longer receives updates or commercial support, and migrating to Alloy is recommended, with a conversion tool [22]. Whoever still has it in production operates without the vendor's backing.

An agent that reads the log, or instrumenting the application?

It is the most practical decision in the whole topic, and it is framed poorly when presented as "the old against the modern." They are two paths with different advantages.

  • An agent that reads is a program that takes what the software already writes —to standard output or to a file— and carries it to the store. It doesn't require changing the software. In exchange, it requires interpreting (parsing) lines in diverse formats, and it cannot invent a trace_id that the application didn't write [32][40]. OpenTelemetry describes the two paths —reading files or standard output, or sending directly over OTLP— and sums up the trade-off like this: the first requires no changes but demands robust parsing; the second avoids the complexity of files but requires configuration in the application [39].
  • Instrumenting from the inside means using the OpenTelemetry SDK or an appender —an adapter that connects the logging system the framework already uses, such as Logback or Log4j in Java, with OpenTelemetry— so that each record comes out with its trace context and structured data [38]. OpenTelemetry doesn't replace those logging libraries: it bridges them [38]. In exchange, you have to modify and deploy code.
An agent that reads the log, or instrumenting the application?
SituationBest fitWhy
Legacy, third-party or certified software that must not be recompiledAgent that readsCoverage without recompiling, once the agent is deployed and the format is interpreted; it is enough for the software to write to standard output
Databases, load balancers, proxies and platform componentsAgent that readsThey rarely incorporate an application SDK
Critical business flow whose exact path is neededInstrument (SDK and appender)It can add the active context and business data before the record leaves the process [38]
Your own Java application with Logback or Log4jInstrument with an appender or with context in the logThe Java agent can inject trace_id and span_id into the context of each line [38]
Many languages and teams at onceAgent that reads as a common safety netA single mechanism for everyone, while instrumenting in stages
High volume of low-value debug messagesAgent with filtersIt discards noise before it reaches the store and costs money
A language where OpenTelemetry logs are still in development [5]Agent that reads, with the trace_id written by the applicationIt avoids relying on a component that is not yet stable

The mature answer is usually hybrid, with one condition: don't duplicate. If the application exports an event through its appender and, in addition, that same event is read again from the file, the store receives two copies. Each flow is declared for one path or the other.

A note on "automatic instrumentation." In Java, the OpenTelemetry agent is attached at startup (-javaagent) and injects code at runtime to capture telemetry from common libraries: incoming and outgoing HTTP calls, the database [37]. It gives quick visibility without changing the source code. But it covers the technical perimeter, not the business logic: the documentation is explicit that an application's own code is normally not instrumented on its own [36]. A decision such as "credit approval" or "inventory allocation" only appears in the trace if someone marked it in the code. What automatic instrumentation buys is a good starting point, not the end of the work.

What to look at in each industry

These are illustrative scenarios, with no customers or measured figures. Each one walks through the same chain: metric that warns, trace that narrows down the path, log that supplies the detail.

Banking: transfers and authorizations. In metrics, you watch the time within which 95% of transfers respond and the authorization error rate. When it rises, the exemplar opens the trace of a slow transfer; the trace shows that the time went into the query to the core system, and the log with the same trace_id says that the the connection pool timed out. With that chain set up and tested, knowing where and why stops depending on gathering five teams; how long it takes depends on the cross-referencing being configured as described above. Warning: these metrics are for operating, not for reconciling [9].

Fintech: open APIs. A partner that consumes your APIs reports that payments "sometimes" fail. Without telemetry it is an anecdote. With traces you can search only for the failed operations of that route and see whether they coincide with the slowness of an external antifraud query. You also watch the error rate per partner, the rejections due to request limits and the latency of notifications (webhooks). That is how "our system failed" is separated from "a third party degraded," with evidence.

Insurance: quoting and issuing. A quote takes 40 seconds at peak hour. The trace reveals that each quote queries three rating systems in series, and the error metric shows that one of them degrades at noon. The decision is to parallelize or add a cache, not to "buy more servers." In issuing, you measure end-to-end time and where the queues build up.

Retail: checkout and inventory. During a promotion, the cart "keeps thinking." The metrics show saturation of the inventory service; the trace reveals that each visit to a product queries it twelve times. The cause is fixed, not the symptom. And here a concrete cost appears: if metrics are labeled by product or by customer, cardinality shoots up (further below).

Regulated company: invoicing. An auditor asks what happened with an invoicing request on Tuesday. With the trace identifier, the path through all the systems is reconstructed, and in the associated logs you can see what was done and when. With one caveat: telemetry helps to understand technical behavior, but it does not replace a formal, intact audit record with approved retention. Traces are sampled and kept for a short time; that is decided with Risk, Privacy and Legal, not with a technical configuration.

Costs and risks: what is rarely said

In many cases, the relevant cost comes from volume, cardinality, retention, query compute and operation; the weight of each item depends on the platform and the contract. Three decisions weigh especially.

Cardinality: every new label is a bill

Cardinality is the number of distinct combinations of labels. In Prometheus, each new combination is a new series that consumes memory, disk and compute. The project's guide recommends keeping the cardinality of each metric below 10 and, if it goes beyond about 100 or can grow without bound, looking for another solution; its example is clear: 10,000 nodes with dozens of filesystems give around 100,000 series, acceptable, but adding a per-user quota produces millions [10]. Labels are not used for user identifiers or emails either [11]. These figures are a project best-practices guide, not a hard technical limit.

Loki has its own version: labels with unbounded values force building an enormous index and thousands of tiny chunks, and the system performs very poorly. Grafana's guide —a vendor figure— suggests not exceeding 10 to 15 labels and moving identifiers to the content or to structured metadata [14].

Translated into business terms: in the year-end promotion of the retail scenario, if each customer is a label, the bill and the slowness grow exactly when the system is needed most.

Personal data: what is not emitted cannot leak

Logs and traces have the bad habit of capturing too much. A service that, by mistake, writes the full name and the identity document of whoever uploads a file leaves that text in the store, with its retention, available to whoever has access to the dashboard. OpenTelemetry is clear: it cannot know what is sensitive in your context, regulatory responsibility lies with whoever implements, and avoiding emitting the data is better than remediating it afterwards; it even warns that hashing an identifier may not provide anonymity when the universe of values is small [42].

The Collector offers processors to remove or transform attributes, but their maturity should be measured: for logs, redaction and filter are listed as alpha and transform as beta [43]. Resting all the protection on an alpha component is fragile. The prudent practice is two barriers: don't emit the data from the application and, as a second line, filter in the Collector. In Mexico, the Federal Law on the Protection of Personal Data Held by Private Parties currently in force (published in the Official Gazette on March 20, 2025, which repealed the 2010 one) aims to regulate legitimate, controlled and informed processing, and requires the data controller to maintain administrative, technical and physical security measures [45]; what is kept and for how long is validated with Privacy and Legal.

Retention and disk: nobody decides it for you

Kubernetes doesn't retain logs long term, and Docker, by default, can fill the disk [27][28]. Long-term retention is a decision for the store and for the business, and it must distinguish between metrics, logs, traces and regulatory evidence. We don't include retention figures or prices because they depend on the vendor and the contract; any that is cited without that information is an assumption.

Who sees what, and who operates it. Grafana dashboards show what the logs and traces contain, so access is designed by area: each team sees the dashboards and sources that correspond to it, and whoever queries logs with personal data is a smaller group than whoever looks at a latency chart. And operating this set of tools requires at least one person who understands the queries (PromQL, LogQL, TraceQL) and the Collector; it is a human cost, not just an infrastructure one, that is worth counting from the start.

On top of that comes sampling: keeping all traces is expensive, and keeping only some forces you to decide which, for example all the ones that fail and a fraction of the healthy ones. With sampling done at the gateway, the routing has to send all the spans of the same trace to the same collector [8]. And a note on how the protocol delivers: on transient failures an implementation may retry sending; if no acknowledgment was received, that can produce duplicates. Storage and queries must therefore tolerate them, without assuming guaranteed delivery [4].

How to start without trying to do everything

  1. Choose a critical business flow, not "the whole system." A transfer, a quote, a checkout, an invoice. If it fails, it hurts; that is why it is instrumented first.
  2. Define in advance the three or four questions you want to be able to answer and the signal that answers them: what does the metric measure, what does the trace mark, what should the log say?
  3. Agree on names from the start. A consistent service name and environment across the three signals are worth more than any dashboard [46].
  4. Start with what doesn't require touching the software. An agent that reads standard output and, in Java, automatic instrumentation give a first view in a short time [36][37].
  5. Instrument from the inside what matters to the business: the steps of that flow, with the trace_id on every log line.
  6. Set cardinality and personal data rules before turning on the tap: a list of permitted attributes, which data is not emitted and who decides retention.
  7. Test the cross-referencing end to end: from the click on the chart to the trace and from the trace to the log. If any jump fails, that is when you learn what was missing, not in a real incident.
  8. Retire what no longer has support. If you have Promtail, plan the migration [22].

Before starting, it is worth agreeing in writing, for the first flow: which flow is included, who owns it, the baseline and the latency and error target, the minimum trace coverage, the prohibited attributes, retention, the monthly budget and the correlation test from metric to trace and to log.

Knowing which tool exists is the easy part. The hard part is knowing which part of your platform already emits what is needed, what can be observed without touching the software and what requires instrumentation, where personal data is slipping into the logs today and what cost each decision brings. That is what we review in an assessment: we deliver to you a map of signals for your critical flows, the table of what is covered with an agent and what requires instrumenting, the inventory of personal data that reaches the logs today, the cost factors that weigh most in your case and the recommended order for adopting it, so that IT, security, risk and business decide with the same information.

References

  1. OpenTelemetry, What is OpenTelemetry? (vendor-neutral; "not a backend"). https://opentelemetry.io/docs/what-is-opentelemetry/
  2. OpenTelemetry, Signals and Context propagation (W3C Trace Context, traceparent header). https://opentelemetry.io/docs/concepts/signals/ · https://opentelemetry.io/docs/concepts/context-propagation/
  3. OpenTelemetry, Specification status. https://opentelemetry.io/docs/specs/status/
  4. OpenTelemetry, OTLP specification (v1.11; ports 4317 and 4318; retries and possible duplicates). https://opentelemetry.io/docs/specs/otlp/
  5. OpenTelemetry, status by language: Java, Go, JavaScript and Python (consulted on Oct 6, 2026). https://opentelemetry.io/docs/languages/ · https://opentelemetry.io/docs/languages/java/ · https://opentelemetry.io/docs/languages/go/ · https://opentelemetry.io/docs/languages/js/ · https://opentelemetry.io/docs/languages/python/
  6. OpenTelemetry, Collector. https://opentelemetry.io/docs/collector/
  7. OpenTelemetry, Collector: agent pattern. https://opentelemetry.io/docs/collector/deploy/agent/
  8. OpenTelemetry, Collector: gateway pattern and agent to gateway. https://opentelemetry.io/docs/collector/deploy/gateway/ · https://opentelemetry.io/docs/collector/deploy/other/agent-to-gateway/
  9. Prometheus, Overview (pull model; accuracy and billing). https://prometheus.io/docs/introduction/overview/
  10. Prometheus, Instrumentation (cardinality guide). https://prometheus.io/docs/practices/instrumentation/
  11. Prometheus, Metric and label naming. https://prometheus.io/docs/practices/naming/
  12. Prometheus, Feature flags (exemplar storage, experimental). https://prometheus.io/docs/prometheus/latest/feature_flags/
  13. Prometheus, Using Prometheus as your OpenTelemetry backend. https://prometheus.io/docs/guides/opentelemetry/
  14. Grafana Labs (vendor), Loki: labels. https://grafana.com/docs/loki/latest/get-started/labels/
  15. Grafana Labs (vendor), Loki: overview. https://grafana.com/docs/loki/latest/get-started/overview/
  16. Grafana Labs (vendor), Loki: sending data with OpenTelemetry. https://grafana.com/docs/loki/latest/send-data/otel/
  17. Grafana Labs (vendor), Tempo: metrics from traces. https://grafana.com/docs/tempo/latest/metrics-from-traces/
  18. Grafana Labs (vendor), Tempo: configuration and TraceQL. https://grafana.com/docs/tempo/latest/configuration/ · https://grafana.com/docs/tempo/latest/traceql/
  19. Grafana Labs (vendor), Grafana fundamentals. https://grafana.com/docs/grafana/latest/fundamentals/
  20. Grafana Labs (vendor), Exemplars and Configure the Prometheus data source. https://grafana.com/docs/grafana/latest/fundamentals/exemplars/ · https://grafana.com/docs/grafana/latest/datasources/prometheus/configure/
  21. Grafana Labs (vendor), Configure the Tempo data source (trace to logs, trace to metrics). https://grafana.com/docs/grafana/latest/datasources/tempo/configure-tempo-data-source/
  22. Grafana Labs (vendor), Promtail: end of life and migration to Alloy. https://grafana.com/docs/loki/latest/send-data/promtail/ · https://grafana.com/docs/alloy/latest/set-up/migrate/from-promtail/
  23. Grafana Labs (vendor), Grafana Alloy. https://grafana.com/docs/alloy/latest/ · https://grafana.com/docs/alloy/latest/introduction/
  24. Grafana Labs (vendor), Alloy: logs in Kubernetes. https://grafana.com/docs/alloy/latest/collect/logs-in-kubernetes/
  25. Grafana Labs (vendor), Alloy: deployment. https://grafana.com/docs/alloy/latest/set-up/deploy/
  26. Grafana Labs (vendor), Alloy: from OpenTelemetry to the LGTM stack. https://grafana.com/docs/alloy/latest/collect/opentelemetry-to-lgtm-stack/
  27. Kubernetes, Logging architecture. https://kubernetes.io/docs/concepts/cluster-administration/logging/
  28. Docker, Logging, Configure logging drivers, json-file and local drivers, Compose and Collector in Docker (OpenTelemetry). https://docs.docker.com/engine/logging/ · https://docs.docker.com/engine/logging/configure/ · https://docs.docker.com/engine/logging/drivers/json-file/ · https://docs.docker.com/engine/logging/drivers/local/ · https://docs.docker.com/compose/ · https://opentelemetry.io/docs/collector/install/docker/
  29. Docker, fluentd logging driver. https://docs.docker.com/engine/logging/drivers/fluentd/
  30. OpenTelemetry, Collector components on Kubernetes. https://opentelemetry.io/docs/platforms/kubernetes/collector/components/
  31. OpenTelemetry, Collector Helm chart. https://opentelemetry.io/docs/platforms/kubernetes/helm/collector/
  32. OpenTelemetry Collector Contrib, filelog receiver (README). https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/receiver/filelogreceiver/README.md
  33. OpenTelemetry Collector Contrib, k8sattributes processor (README). https://github.com/open-telemetry/opentelemetry-collector-contrib/blob/main/processor/k8sattributesprocessor/README.md
  34. OpenTelemetry, Operator for Kubernetes, auto-instrumentation and its troubleshooting guide. https://opentelemetry.io/docs/platforms/kubernetes/operator/ · https://opentelemetry.io/docs/platforms/kubernetes/operator/automatic/ · https://opentelemetry.io/docs/platforms/kubernetes/operator/troubleshooting/automatic/
  35. OpenTelemetry, Collector resiliency. https://opentelemetry.io/docs/collector/resiliency/
  36. OpenTelemetry, Zero-code instrumentation. https://opentelemetry.io/docs/concepts/instrumentation/zero-code/
  37. OpenTelemetry, Java agent. https://opentelemetry.io/docs/zero-code/java/agent/
  38. OpenTelemetry, Java instrumentation (Logback and Log4j appenders; trace context in logs). https://opentelemetry.io/docs/languages/java/instrumentation/
  39. OpenTelemetry, Logs specification. https://opentelemetry.io/docs/specs/otel/logs/
  40. OpenTelemetry, Logs data model. https://opentelemetry.io/docs/specs/otel/logs/data-model/
  41. OpenTelemetry, Trace context in non-OTLP log formats. https://opentelemetry.io/docs/specs/otel/compatibility/logging_trace_context/
  42. OpenTelemetry, Handling sensitive data. https://opentelemetry.io/docs/security/handling-sensitive-data/
  43. OpenTelemetry, Collector processors (stability per component). https://opentelemetry.io/docs/collector/components/processor/
  44. CNCF, Annual Survey 2024 (published Apr 1, 2025; community survey, 689 responses on the question cited). https://www.cncf.io/reports/cncf-annual-survey-2024/
  45. Cámara de Diputados, Ley Federal de Protección de Datos Personales en Posesión de los Particulares (new law published in the DOF on Mar 20, 2025; text in force, last amended DOF Nov 14, 2025; articles 1 and 18). https://www.diputados.gob.mx/LeyesBiblio/pdf/LFPDPPP.pdf
  46. OpenTelemetry, Semantic conventions. https://opentelemetry.io/docs/concepts/semantic-conventions/
  47. Grafana Labs (vendor), Configure trace to logs (Loki derived fields toward Tempo). https://grafana.com/docs/grafana/latest/datasources/tempo/configure-tempo-data-source/configure-trace-to-logs/