关于 Distributed Tracing
Alauda Distributed Tracing 记录单个请求在分布式应用中跨服务的运行路径。它帮助团队端到端地跟踪请求路径,理解跨服务边界的延迟,并在云原生微服务环境中排查故障。
在 microservices 架构中,一个用户操作通常会触发多个 service、database、message queue 和基础设施组件上的工作。Distributed Tracing 会将这些工作单元关联为单个 trace,使运维人员可以分析完整的执行路径,而不是孤立地检查每个 service。
trace 表示请求在系统中的完整路径。每个 trace 包含一个或多个 span,每个 span 记录一个逻辑工作单元,包括 operation name、start time、duration 和相关元数据。父子 span 关系展示上游和下游操作之间的连接方式。
Alauda Distributed Tracing 提供以下能力:
- 端到端请求可见性:用于查看分布式事务的完整执行路径
- 延迟分析:用于识别缓慢的 service、操作或下游依赖
- 根因分析:用于定位跨 service 边界的故障
- 可扩展的 trace 存储和查询:基于 Jaeger,并支持多个后端
Alauda Distributed Tracing 可与 Alauda Build of OpenTelemetry 配合使用,通过 OTLP 等 open telemetry 标准接收、处理并转发 trace 数据。它还可以与 Alauda Service Mesh 集成,从而在排查问题时将 service-to-service 流量与 trace 数据一并分析。
Note
因为 Alauda Distributed Tracing 的发版周期与灵雀云容器平台不同,所以 Alauda Distributed Tracing 的文档现在作为独立的文档站点托管在 Alauda Distributed Tracing。