进阶 pytorch.org 2026-10-07 23:07:35 · 6 阅读

第4章 CUDA Graph 内核标注与性能分析实战

注意

前往文末下载完整示例代码。

CUDA Graph 内核标注与性能分析

作者:Shangdi Yu

你将学到什么

  • 如何在内核标注加持下捕获 CUDA graphs
  • 如何对带标注的 graphs 进行性能分析(profiling)
  • 如何导出包含内联标注和语义化内核泳道(semantic kernel lanes)的 trace
  • 如何通过自定义流分配来可视化图执行过程
  • 如何使用 mark_stream 恢复逻辑流泳道
  • 如何为通信集合操作(communication collectives)添加元数据标注(集合类型、消息大小、组、rank),补全 Eager 模式下 NCCL trace 有而 CUDA graph 模式缺失的信息

前置条件

  • PyTorch 2.15+(在 2.15 正式发布前请使用 nightly 版本)
  • 支持 CUDA 的 GPU
  • Driver/CUDA-compat >= 13.1 以支持标注功能
  • cuda-bindings >= 13.1.0

CUDA graphs 是一种强大的优化技术,通过捕获并重放一系列 CUDA 操作序列,能显著降低内核启动开销。然而,在对 CUDA graphs 进行性能分析时,所有内核都显示在同一个流上,这使得理解计算的逻辑结构变得困难。本教程演示如何使用内核标注为 CUDA graphs 中的内核添加语义标签。性能分析器在导出 trace 时可直接包含这些标注,并创建自定义可视化泳道,从而更轻松理解与调试复杂的图执行过程。

标注不仅限于计算内核。最有价值的用法之一是为通信集合操作添加标注。在 Eager 模式下,性能分析器会为每个 NCCL 内核附加丰富元数据——包括集合类型、消息大小、进程组及 rank——让你能看清每个通信操作的具体行为。而在 CUDA graph 模式下,这些元数据会丢失:集合操作作为不透明内核被重放。本教程展示如何通过标注重新附加这些元数据,使带图的通信操作读取起来与 Eager 模式无异。

概述

CUDA graph 内核标注允许你在图捕获期间为内核添加语义标签。这些标签有助于在性能分析时理解每个内核的作用,让你能轻松识别模型中正在执行的特定部分(如 Attention、MLP、Normalization)。

没有标注时,性能分析器 trace 显示所有内核在单一泳道(stream)上,并带有自动生成的名称,导致难以理解计算的逻辑结构。有了标注,你可以:

  • 在捕获时为内核组添加有意义的名称
  • 分配自定义流 ID 以便可视化组织
  • 在性能分析器 trace 中直接导出标签用于语义可视化

结果是,性能分析器 trace 中的内核将根据其功能进行标记和组织,使识别性能瓶颈和理解执行流程变得更加容易。

标注前:所有内核显示在单一泳道,带有自动生成的名称,难以分辨哪些操作属于模型的哪个逻辑组件。

标注后:内核被组织到语义化泳道(如 stream 61 和 62),并带有“attention”、“mlp”等有意义标签,便于识别不同组件并理解执行结构。

另一个示例:这是一个带有标注元数据的 AllReduce 内核:

需求

本教程需要:

  • PyTorch 2.15+(在 2.15 正式发布前请使用 nightly 版本)
  • CUDA GPU
  • Driver/CUDA-compat >= 13.1 以支持标注功能
  • cuda-bindings 包版本 >= 13.1.0 (pip install "cuda-bindings>=13.1.0")

cuda-bindings 包提供 CUDA Runtime API 的 Python 绑定。版本 13.1.0+ 是必需的,因为它支持启用内核标注的 cudaGraphNodeGetToolsId API。如果版本较旧,教程仍可运行,但标注功能将禁用,并会出现提示升级的警告信息。

在旧版 Driver 或 cuda-bindings 上,捕获和性能分析仍可工作,但 mark_kernels 将变为空操作(no-op),最终 trace 中也不会出现语义化泳道。

import gzip
import json
import math
import os
from collections import Counter
from pathlib import Path

import torch
import torch.distributed as dist
import torch.multiprocessing
from torch.profiler import profile, ProfilerActivity
from torch.cuda.graph_annotations import get_kernel_annotations, is_available, mark_kernels
from torch.cuda._graph_annotations import get_stream_for_pg, mark_stream

构建模型

让我们创建一个简单的 Transformer block 作为示例模型。我们将为计算的不同部分(QKV 投影、Attention、输出投影、MLP)添加标注,以便在性能分析器中将其作为独立泳道查看。

def build_transformer_block():
    """创建带有参数的简单 Transformer block。"""
    device = "cuda"
    torch.manual_seed(0)

    # 模型维度
    batch_size, seq_len, dim, num_heads = 4, 256, 1024, 8
    head_dim = dim // num_heads

    # 初始化参数
    params = {
        "x": torch.randn(batch_size, seq_len, dim, device=device),
        "Wqkv": torch.randn(dim, 3 * dim, device=device) / math.sqrt(dim),
        "Wo": torch.randn(dim, dim, device=device) / math.sqrt(dim),
        "W1": torch.randn(dim, 4 * dim, device=device) / math.sqrt(dim),
        "W2": torch.randn(4 * dim, dim, device=device) / math.sqrt(4 * dim),
    }

    def forward():
        """带有标注区域的前向传播。"""
        B, T, D, H = batch_size, seq_len, dim, num_heads
        hd = head_dim

        # 标注 QKV 投影
        with mark_kernels({"name": "qkv_proj"}):
            qkv = params["x"] @ params["Wqkv"]

        # 重塑以适配多头注意力
        q, k, v = qkv.split(D, dim=-1)
        q = q.view(B, T, H, hd).transpose(1, 2)
        k = k.view(B, T, H, hd).transpose(1, 2)
        v = v.view(B, T, H, hd).transpose(1, 2)

        # 标注注意力计算(可选使用自定义流)
        with mark_kernels({"name": "attention", "stream": 62}):
            scores = (q @ k.transpose(-1, -2)) / math.sqrt(hd)
            attn = torch.softmax(scores, dim=-1)
            ctx = (attn @ v).transpose(1, 2).reshape(B, T, D)

        # 标注输出投影
        with mark_kernels({"name": "out_proj"}):
            o = ctx @ params["Wo"]

        # 标注 MLP(使用另一个自定义流)
        with mark_kernels({"name": "mlp", "stream": 61}):
            return torch.nn.functional.gelu(o @ params["W1"]) @ params["W2"]

    return forward

mark_kernels 上下文管理器

核心 API 是 mark_kernels(),它接受一个字典,包含:

  • name:存储在内核事件 args 中的字符串标签(当分配自定义流时,也用作泳道名称)
  • stream(可选):用于可视化的虚拟流 ID

在该上下文内启动的任何 CUDA 内核都会被标记这些标注。当我们使用 graph_lanes="all" 导出性能分析器 trace 时,这些标签会将内核组织到自定义泳道中。流 ID 仅控制可视化;不会改变内核的实际执行位置。

捕获带标注的 CUDA Graph

要捕获启用标注的图,需向 torch.cuda.graph() 传递 enable_annotations=True。这将自动处理标注的生命周期:启用、解析和重映射。

def capture_graph_with_annotations(model_fn):
    """将模型捕获为启用标注的 CUDA graph。"""
    # 在侧流(side stream)上进行预热
    warmup_stream = torch.cuda.Stream()
    warmup_stream.wait_stream(torch.cuda.current_stream())

    with torch.cuda.stream(warmup_stream):
        for _ in range(3):
            model_fn()

    torch.cuda.current_stream().wait_stream(warmup_stream)

    # 启用标注进行捕获
    graph = torch.cuda.CUDAGraph()
    with torch.cuda.graph(graph, enable_annotations=True):
        output = model_fn()

    num_annotations = len(get_kernel_annotations())
    print(f"Captured graph with {num_annotations} annotated nodes")

    return graph, output

对图进行性能分析

捕获图后,先重放几次以预热,然后对后续的重放进行性能分析。将记录的标注直接传递给 export_chrome_trace(),以将其包含在导出的内核事件中。我们还从同一性能分析会话导出原始 trace 以便对比。

当映射非空时,cuda_graph_annotations 会自动选择 Python 导出器。此处显式选择它以确保即使映射为空也能工作。两个文件都可以在 https://ui.perfetto.dev/ 中打开。

def profile_graph(graph, output_dir):
    """对图重放进行性能分析,并导出原始和带标注的 trace。"""
    output_dir = Path(output_dir)
    output_dir.mkdir(exist_ok=True, parents=True)

    # 预热重放
    for _ in range(3):
        graph.replay()
    torch.cuda.synchronize()

    # 对多次重放进行性能分析
    with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA]) as prof:
        for _ in range(5):
            graph.replay()
        torch.cuda.synchronize()

    raw_trace_path = output_dir / "trace_raw.json.gz"
    prof.export_chrome_trace(str(raw_trace_path), use_python_export=True)
    print(f"Saved raw trace to {raw_trace_path}")

    annotations = get_kernel_annotations()
    annotated_path = output_dir / "trace_annotated.json.gz"
    prof.export_chrome_trace(
        str(annotated_path),
        use_python_export=True,
        cuda_graph_annotations=annotations,
        graph_lanes="all" if annotations else "none",
    )
    print(f"Saved annotated trace to {annotated_path}")

    return raw_trace_path, annotated_path

选择 Trace 布局

graph_lanes="all" 将带有流标注的图事件放置在对应的显示泳道上。其他图事件去 default_stream(默认为 7)。被移动的事件会在 args["original_stream"] 中保留其实际执行流。在我们的示例中,这将 Attention 分组在泳道 62,MLP 在泳道 61。

若要保留记录的流布局,请省略 graph_lanes 或设为 "none"。标注元数据仍包含在每个匹配事件的 args 中,任何被标注的流都存储为 args["annotated_stream"]。这在检查原始流上的并发时非常有用。

空的标注映射被视为无标注。由于 graph_lanes="all" 需要非空映射,因此当标注支持不可用时,示例会使用 "none"。

请保持捕获的图活跃直到导出完成:销毁或重置图会将其条目从实时标注注册表中移除。无需将标注保存到单独文件或对导出的 trace 进行后处理。

对比前后效果

为了查看标注的影响,让我们统计内核在线程 ID(代表 trace 中的可视化泳道)上的分布情况。

def compare_traces(raw_trace, annotated_trace):
    """对比标注前后内核的分布。"""
    def count_lanes(trace):
        """统计每个泳道(tid)的内核数量。"""
        counter = Counter(
            event["tid"]
            for event in trace["traceEvents"]
            if event.get("cat") == "kernel"
        )
        return dict(sorted(counter.items()))

    raw_lanes = count_lanes(raw_trace)
    annotated_lanes = count_lanes(annotated_trace)

    print("\n" + "="*60)
    print("BEFORE annotation - kernels per lane (tid -> count):")
    for tid, count in raw_lanes.items():
        print(f"  Stream {tid}: {count} kernels")

    print("\nAFTER annotation - kernels per lane (tid -> count):")
    for tid, count in annotated_lanes.items():
        print(f"  Stream {tid}: {count} kernels")
    print("="*60)

完整流程

现在让我们运行完整的工作流:构建模型、带标注捕获、性能分析以及导出带标注的 trace。

def main():
    """端到端 CUDA graph 标注与性能分析演示。"""
    if not torch.cuda.is_available():
        raise SystemExit("CUDA required for this tutorial")

    # 检查标注支持是否可用
    # 如果 cuda-bindings 版本过旧,PyTorch 会记录警告
    supported = is_available()
    print(f"Annotation support available: {supported}")
    if not supported:
        print("NOTE: Annotation API not available.")
        print("This could be due to:")
        print("  - Driver/CUDA-compat < 13.1")
        print("  - Outdated cuda-bindings (check PyTorch warnings above)")
        print("Annotations will not be recorded, but the demo will still run.")
        print("The exported trace will keep the recorded stream layout.\n")

    output_dir = Path("traces")

    # 构建模型
    print("\n1. Building transformer block model...")
    model_fn = build_transformer_block()

    # 捕获带标注的图
    print("\n2. Capturing CUDA graph with annotations...")
    graph, output = capture_graph_with_annotations(model_fn)

    print("\n3. Profiling graph replays and exporting traces...")
    raw_trace_path, annotated_path = profile_graph(graph, output_dir)

    print("\n4. Comparing traces...")
    with gzip.open(raw_trace_path, "rt") as f:
        raw_trace = json.load(f)
    with gzip.open(annotated_path, "rt") as f:
        annotated_trace = json.load(f)
    compare_traces(raw_trace, annotated_trace)

    # 总结
    print("\n" + "="*60)
    print("SUMMARY")
    print("="*60)
    print(f"Raw trace:       {raw_trace_path}")
    print(f"Annotated trace: {annotated_path}")
    print("\nOpen the annotated trace in https://ui.perfetto.dev/ to visualize")
    print("the semantic kernel lanes.")
    print("="*60)

# 示例输出:
# if __name__ == "__main__":
#     main()
#
# Annotation support available: True
#
# 1. Building transformer block model...
#
# 2. Capturing CUDA graph with annotations...
# Captured graph with 13 annotated nodes
#
# 3. Profiling graph replays and exporting traces...
# Saved raw trace to traces/trace_raw.json.gz
# Saved annotated trace to traces/trace_annotated.json.gz
#
# 4. Comparing traces...
#
# ============================================================
# BEFORE annotation - kernels per lane (tid -> count):
#   Stream 7: 65 kernels
#
# AFTER annotation - kernels per lane (tid -> count):
#   Stream 7: 10 kernels
#   Stream 61: 15 kernels
#   Stream 62: 40 kernels
# ============================================================
#
# ============================================================
# SUMMARY
# ============================================================
# Raw trace:       traces/trace_raw.json.gz
# Annotated trace: traces/trace_annotated.json.gz
#
# Open the annotated trace in https://ui.perfetto.dev/ to visualize
# the semantic kernel lanes.
# ============================================================

使用 mark_stream 记录逻辑流

如果你的计算已经使用了多个 CUDA 流,mark_stream 会切换到该流并为其内核记录一个逻辑泳道 ID。此辅助函数目前位于私有模块 torch.cuda._graph_annotations 中。如果你只想更改显示布局(如上面的 Transformer 示例),请在使用 mark_kernels 时显式指定 stream 字段。

下面的代码块在两个独立的 CUDA 流上运行两个独立的投影,然后合并它们的结果。每个侧流在其 mark_stream 范围内等待当前流。这些等待操作使侧流在启动工作前加入捕获。当前流在读取输出和结束捕获前等待这两个分支。mark_stream 不会自动为你插入这些依赖关系。

def build_multistream_block():
    """在两个不同的带标注 CUDA 流上创建两个投影。"""
    x = torch.randn(1024, 1024, device="cuda")
    left_weight = torch.randn_like(x)
    right_weight = torch.randn_like(x)
    left = torch.empty_like(x)
    right = torch.empty_like(x)
    left_stream = torch.cuda.Stream()
    right_stream = torch.cuda.Stream()

    def forward():
        current_stream = torch.cuda.current_stream()
        with mark_stream(left_stream, "left_projection"):
            left_stream.wait_stream(current_stream)
            torch.mm(x, left_weight, out=left)

        with mark_stream(right_stream, "right_projection"):
            right_stream.wait_stream(current_stream)
            torch.mm(x, right_weight, out=right)

        current_stream.wait_stream(left_stream)
        current_stream.wait_stream(right_stream)
        with mark_kernels("combine"):
            return left + right

    return forward

复用捕获和性能分析辅助函数来导出此图。使用 graph_lanes="all" 时,即使图重放将它们调度在不同的硬件流上,投影也会出现在名为 left_projection 和 right_projection 的独立逻辑泳道上。Combine 内核去往默认泳道 7。被移动的事件保留 args["original_stream"]。泳道 ID 是自动分配的,可能因已标记的流不同而变化。重用相同流会重用其泳道 ID。将当前流传递给 mark_stream 会添加标签而不分配新泳道。使用 graph_lanes="none" 时,保留原始布局,记录的泳道 ID 出现在 args["annotated_stream"] 中。

def stream_annotation_demo():
    """捕获并分析两个投影的逻辑流泳道。"""
    model_fn = build_multistream_block()
    graph, output = capture_graph_with_annotations(model_fn)
    _, annotated_path = profile_graph(graph, "traces_streams")
    return annotated_path

运行 stream_annotation_demo() 并在 https://ui.perfetto.dev/ 中打开 traces_streams/trace_annotated.json.gz。

标注通信集合操作

在 Eager 模式下,性能分析器自动拦截 NCCL 集合操作并记录丰富元数据:集合类型、输入/输出消息大小、进程组、其大小以及参与 rank。

在 CUDA graph 中,这种自动拦截停止工作。集合操作被捕获一次,然后作为不透明内核节点重放。性能分析器无法拦截图重放,因此没有地方附加 NCCL 元数据。内核仍然显示在 trace 中(例如 ncclDevKernel_AllReduce_Sum_f32_RING_LL),但它们是不透明的:你无法得知集合类型、移动了多少字节或属于哪个进程组。

标注填补了这一空白。通过在 mark_kernels 中包装集合操作,并使用性能分析器在 Eager 模式下自动附加的相同字段,我们可以手动将该元数据重新附加到带图的内核上。导出后,带图的集合操作读取起来与 Eager 模式无异。下面的辅助函数构建元数据字典;使用性能分析器在 Eager 模式下使用的字段名(如 In msg nelems、Group size、Process Group Name、等)可保持带标注的 trace 与非带图 trace 一致。

def annotate_collective(collective_name, input_tensor, output_tensor, group=None):
    """为集合操作添加 Eager NCCL trace 中暴露的元数据。

    返回一个 ``mark_kernels`` 上下文管理器。在内核(即集合操作)内部启动的内核
    将被标记集合类型、消息大小、dtype、以及进程组的名称/描述/rank,
    并放置在以进程组为键的专用泳道上,以便在视觉上与计算分开。

    字段名与性能分析器为 Eager 集合操作记录的键匹配
    (``In msg nelems``、``Group size``、``Process Group Name``...),
    因此带标注的带图集合操作读取起来与非带图集合操作完全一致。
    """
    pg = group if group is not None else (dist.group.WORLD if dist.is_initialized() else None)
    ranks = dist.get_process_group_ranks(pg) if pg is not None else [0]
    group_name = getattr(pg, "group_name", "default")
    group_desc = getattr(pg, "group_desc", "default")

    # NCCL 始终使用其内部流,因此以进程组(名称 + 描述)为键分配泳道,
    # 并给予其稳定 ID(>= 60)。
    pg_key = f"{group_name}_{group_desc}"
    annotation = {
        "name": collective_name,
        "In msg nelems": input_tensor.numel(),
        "Out msg nelems": output_tensor.numel(),
        "Group size": len(ranks),
        "dtype": str(input_tensor.dtype).replace("torch.", ""),
        "Process Group Name": group_name,
        "Process Group Description": group_desc,
        "Process Group Ranks": ranks,
        "stream": get_stream_for_pg(pg_key),
    }
    return mark_kernels(annotation)

混合计算与通信的块

张量并行或数据并行层将矩阵乘法与集合操作交错。在此,投影输出在整个组上进行 All-Reduce,镜像张量并行线性层中的通信。集合操作使用 annotate_collective 进行标注,并落在其专用泳道上。

def build_comm_block(group=None):
    """创建带有标注的计算 + 集合操作块用于性能分析。"""
    device = "cuda"
    torch.manual_seed(0)
    dim = 1024
    params = {
        "x": torch.randn(4, 256, dim, device=device),
        "W": torch.randn(dim, dim, device=device) / math.sqrt(dim),
    }

    def forward():
        with mark_kernels({"name": "proj", "stream": 61}):
            h = params["x"] @ params["W"]

        # 在整个组上 All-Reduce 投影输出(例如张量并行)。
        # all_reduce 是原地操作,因此输入和输出张量相同。
        # 标注重新附加了 CUDA graph 否则会丢弃的 NCCL 元数据。
        if dist.is_available() and dist.is_initialized():
            with annotate_collective("all_reduce", h, h, group):
                dist.all_reduce(h)
        return h

    return forward

运行通信演示

WORLD_SIZE = 2

def init_pg(rank, world_size):
    """为生成的演示中的一个 rank 初始化 NCCL 组。"""
    os.environ["MASTER_ADDR"] = "127.0.0.1"
    os.environ["MASTER_PORT"] = "29500"
    os.environ["RANK"] = str(rank)
    os.environ["WORLD_SIZE"] = str(world_size)
    # 单机设置使用环回接口
    os.environ["NCCL_SOCKET_IFNAME"] = "lo"
    dist.init_process_group("nccl", rank=rank, world_size=world_size)
    torch.cuda.set_device(rank)

def _comm_worker(rank, world_size):
    """每个 rank 的工作者:构建、捕获、分析,并在(rank 0)导出。"""
    init_pg(rank, world_size)

    output_dir = Path("traces_comm")

    if rank == 0:
        print("\nBuilding compute + collective block...")
    model_fn = build_comm_block()

    if rank == 0:
        print("Capturing CUDA graph with annotations...")
    graph, _ = capture_graph_with_annotations(model_fn)

    # 每个 rank 在性能分析期间参与集合操作,但只有 rank 0 导出 trace。
    if rank == 0:
        _, annotated_path = profile_graph(graph, output_dir)
        with gzip.open(annotated_path, "rt") as f:
            annotated_trace = json.load(f)

        # 打印带标注的集合操作内核的 args,以展示 Eager 风格的
        # 元数据现已附加到带图的通信上。
        print("\nAnnotated collective kernels (metadata restored):")
        for event in annotated_trace["traceEvents"]:
            args = event.get("args", {})
            if args.get("In msg nelems") is not None:
                print(f"  {event.get('name', '?')[:40]}")
                for key in (
                    "In msg nelems",
                    "Out msg nelems",
                    "Group size",
                    "dtype",
                    "Process Group Name",
                    "Process Group Description",
                    "Process Group Ranks",
                    "stream",
                ):
                    if key in args:
                        print(f"      {key}: {args[key]}")
        print(f"\nAnnotated trace: {annotated_path}")
    else:
        # 匹配 rank 0 的预热 + 被分析的重放次数,以确保集合操作完成。
        for _ in range(3):
            graph.replay()
        torch.cuda.synchronize()
        for _ in range(5):
            graph.replay()
        torch.cuda.synchronize()

    # 在导出后、销毁组之前释放捕获的 NCCL 工作。
    graph.reset()
    dist.destroy_process_group()

def comm_annotation_demo():
    """生成一个 ``world_size=2`` 组并展示通信元数据。"""
    if not (dist.is_available() and torch.cuda.is_available()):
        print("Distributed/NCCL unavailable; skipping comm annotation demo.")
        return
    if torch.cuda.device_count() < WORLD_SIZE:
        print(f"Need {WORLD_SIZE} GPUs for the comm demo; skipping.")
        return

    torch.multiprocessing.spawn(
        _comm_worker, args=(WORLD_SIZE,), nprocs=WORLD_SIZE, join=True
    )

# 示例输出(2 GPUs):
# if __name__ == "__main__":
#     comm_annotation_demo()
#
# Building compute + collective block...
# Capturing CUDA graph with annotations...
# Captured graph with 3 annotated nodes
# Saved raw trace to traces_comm/trace_raw.json.gz
# Saved annotated trace to traces_comm/trace_annotated.json.gz
#
# The all_reduce runs a real NCCL kernel
# (``ncclDevKernel_AllReduce_Sum_f32_RING_LL``) across the two ranks:
#
# Annotated collective kernels (metadata restored):
#   ncclDevKernel_AllReduce_Sum_f32_RING_LL
#       In msg nelems: 1048576
#       Out msg nelems: 1048576
#       Group size: 2
#       dtype: float32
#       Process Group Name: 0
#       Process Group Description: default_pg
#       Process Group Ranks: [0, 1]
#       stream: 60
#
# In the trace viewer, the all-reduce sits on its own dedicated comm lane
# (stream 60), and selecting it shows the collective type, message sizes, group,
# and ranks -- the same fields you would see in an eager trace, now recovered
# for a CUDA-graphed collective. This metadata is LOST without annotations.

性能考量

内核标注增加的开销极小:

  • 标注标记发生在图捕获期间(一次性成本)
  • 图重放性能与未标注的图相同
  • 标注在性能分析完成后,于 trace 导出阶段添加

主要成本是性能分析本身,这在你优化性能时无论如何都需要做。标注通过添加语义结构,使性能分析器输出更有用。

故障排除

Trace 中没有标注?

  • 检查 Driver/CUDA-compat >= 13.1
  • 验证 torch.cuda.graph() 是否传递了 enable_annotations=True
  • 确保安装了 cuda-bindings>=13.1.0
  • 在图仍活跃时,将 cuda_graph_annotations=get_kernel_annotations() 传递给 export_chrome_trace()

特定内核未显示标注?

  • 某些操作可能不启动内核(例如,张量视图 tensor views)
  • 只有 mark_kernels 上下文中启动的内核会被标注
  • 使用 torch.profiler 验证操作是否实际产生了 CUDA 内核

结论

CUDA graph 内核标注提供了一种强大的方式为性能分析 trace 添加语义结构。通过在图捕获期间标记模型的逻辑组件,并在导出时包含这些标注,你可以创建更容易理解和优化复杂 CUDA graph 执行的可视化。

关键要点:

  • 使用 mark_kernels() 在图捕获期间标记区域
  • 使用 mark_stream() 在切换 CUDA 流时记录逻辑泳道
  • 使用 enable_annotations=True 启用标注
  • 标注通信集合操作以恢复 CUDA graphs 丢弃但 Eager trace 暴露的 NCCL 元数据(集合类型、消息大小、组、rank)
  • 将标注直接传递给 export_chrome_trace()
  • 使用 graph_lanes="all" 将带图内核组织到语义化泳道
  • 在 https://ui.perfetto.dev/ 中查看结果以获得直观的可视化

对于拥有许多组件的大型模型、分布式训练设置或任何理解执行结构对性能优化至关重要的场景,这项技术特别有价值。

脚本总运行时间:(0 分钟 0.007 秒)

下载 Jupyter notebook: cuda_graph_annotations_tutorial.ipynb

下载 Python 源代码: cuda_graph_annotations_tutorial.py

下载压缩包: cuda_graph_annotations_tutorial.zip

评论 (0)