Skip to content
frgfmPublic

About

Class activation maps for your PyTorch models (CAM, Grad-CAM, Grad-CAM++, Smooth Grad-CAM++, Score-CAM, SS-CAM, IS-CAM, XGrad-CAM, Layer-CAM, Finer-CAM, LeGrad, RefineCAM)

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2.3k stars

Watchers

10 watching

Forks

TorchCAM: class activation explorer

English | 简体中文

CI Status ruff ty Test coverage percentage

PyPi Version GitHub release (latest by date) pyversions License

Huggingface Spaces Open in Colab

Documentation Status

Simple way to leverage the class-specific activation of convolutional and transformer layers in PyTorch.

Debugging one surprising classifier result? Use the predicted-versus-expected agent workflow, install the portable skill in your coding agent with npx skills add frgfm/torch-cam, or start from llms.txt.

Source: image from woopets (activation maps created with a pretrained Resnet-18)

Why TorchCAM

  • 12 CAM methods, one API: from CAM and Grad-CAM to Finer-CAM, LeGrad and RefineCAM.
  • CNNs and Vision Transformers: automatic target-layer resolution for CNNs, LeGrad and reshape transforms for ViTs.
  • Lean: fully typed, with only 4 runtime dependencies (PyTorch, NumPy, Pillow, Matplotlib).
  • Measurable: built-in faithfulness metrics (average drop, increase in confidence, deletion/insertion).
  • Agent-ready: explain() saves a manifest-backed evidence bundle that humans and coding agents can verify.

Quick Tour

Explain a prediction

from urllib.request import urlretrieve
from PIL import Image
from torchvision.models import ResNet18_Weights, resnet18
from torchcam.explain import explain

urlretrieve("https://github.com/pytorch/hub/raw/master/images/dog.jpg", "dog.jpg")
image = Image.open("dog.jpg").convert("RGB")
weights = ResNet18_Weights.DEFAULT
model = resnet18(weights=weights).eval()

result = explain(model, weights.transforms()(image).unsqueeze(0), class_names=weights.meta["categories"])
result.save("torchcam-explanation", image)  # CAMs, heatmaps, overlays and manifest.json

Pass expected_class_idx to compare the prediction with the class you expected. See Debug one prediction for the full contract.

Setting your CAM

TorchCAM leverages PyTorch hooking mechanisms to seamlessly retrieve all required information to produce the class activation without additional efforts from the user. Each CAM object acts as a wrapper around your model.

You can find the exhaustive list of supported CAM methods in the documentation, then use it as follows:

from torchvision.models import get_model, get_model_weights
from torchcam.methods import LayerCAM

# Define your model
model = get_model("resnet18", weights=get_model_weights("resnet18").DEFAULT).eval()
# Set your CAM extractor
cam_extractor = LayerCAM(model)

Please note that by default, the layer at which the CAM is retrieved is set to the last non-reduced convolutional layer. If you wish to investigate a specific layer, use the target_layer argument in the constructor.

Retrieving the class activation map

Once your CAM extractor is set, you only need to use your model to infer on your data as usual. If any additional information is required, the extractor will get it for you automatically.

from urllib.request import urlretrieve
from torchvision.io import decode_image
from torchvision.models import get_model, get_model_weights

# Get a model and an image
weights = get_model_weights("resnet18").DEFAULT
model = get_model("resnet18", weights=weights).eval()
preprocess = weights.transforms()
urlretrieve("https://github.com/pytorch/hub/raw/master/images/dog.jpg", "dog.jpg")
img = decode_image("dog.jpg")

input_tensor = preprocess(img)

Compute the class activation map:

from torchcam.methods import LayerCAM

with LayerCAM(model) as cam_extractor:
  out = model(input_tensor.unsqueeze(0))
  # Retrieve the CAM by passing the class index and the model output
  activation_map = cam_extractor(out.squeeze(0).argmax().item(), out)

Here class_idx (the first argument) is the index in the model's output logits of the class you want to explain — out.squeeze(0).argmax().item() picks the top prediction, but you can pass any class index. The extractor returns one activation map per target layer.

If you want to visualize your heatmap, you only need to cast the CAM to a numpy ndarray:

import matplotlib.pyplot as plt
# Visualize the raw CAM
plt.imshow(activation_map[0].squeeze(0).numpy()); plt.axis('off'); plt.tight_layout(); plt.show()

raw_heatmap

Or if you wish to overlay it on your input image:

import matplotlib.pyplot as plt
from torchvision.transforms.v2.functional import to_pil_image
from torchcam.utils import overlay_mask

# Resize the CAM and overlay it
result = overlay_mask(to_pil_image(img), to_pil_image(activation_map[0].squeeze(0), mode='F'), alpha=0.5)
plt.imshow(result); plt.axis('off'); plt.tight_layout(); plt.show()

overlayed_heatmap

Tip

Using your own (non-torchvision) model, a Vision Transformer, 3D/video data, or batched inputs? Read the Advanced usage guide — it also covers how to choose the right target_layer and CAM method.

Hitting a cannot register a hook ... / requires grad error, a NaN, or a blank heatmap? See Troubleshooting.

Setup

Python 3.11 (or higher) and uv/pip are required to install TorchCAM. TorchCAM 0.5.0 requires PyTorch 2.4.1 or higher (torch>=2.4.1,<3) and supports NumPy 1.26.4 and NumPy 2 (numpy>=1.26.4,<3). Use the matching torchvision release.

On macOS, TorchCAM 0.5.0 requires Apple Silicon when using published PyTorch wheels. Intel Mac users can use Python 3.11 or 3.12 and install the previous release with pip install "torchcam==0.4.1".

Stable release

You can install the last stable release of the package using pypi as follows:

pip install torchcam

Latest version

Alternatively, if you wish to use the latest features of the project that haven't made their way to a release yet, you can install the package from source:

pip install "torchcam @ git+https://github.com/frgfm/torch-cam.git"

CAM Zoo

This project is developed and maintained by the repo owner, but the implementation was based on the following research papers:

  • Learning Deep Features for Discriminative Localization: the original CAM paper
  • Grad-CAM: GradCAM paper, generalizing CAM to models without global average pooling.
  • Grad-CAM++: improvement of GradCAM++ for more accurate pixel-level contribution to the activation.
  • Smooth Grad-CAM++: SmoothGrad mechanism coupled with GradCAM.
  • Score-CAM: score-weighting of class activation for better interpretability.
  • SS-CAM: SmoothGrad mechanism coupled with Score-CAM.
  • IS-CAM: integration-based variant of Score-CAM.
  • XGrad-CAM: improved version of Grad-CAM in terms of sensitivity and conservation.
  • Layer-CAM: Grad-CAM alternative leveraging pixel-wise contribution of the gradient to the activation.
  • Finer-CAM: contrastive CAM objective highlighting differences between similar classes.
  • LeGrad: layerwise positive attention-gradient maps for Vision Transformers.
  • RefineCAM: multi-layer refinement producing high-resolution activation maps.

Not sure which one to use? See Choosing a CAM method.

Source: YouTube video (activation maps created by Layer-CAM with a pretrained ResNet-18)

What else

Documentation

The full package documentation is available here for detailed specifications.

Playground app

A minimal demo app is provided for you to play with the supported CAM methods! Feel free to check out the live demo on Hugging Face Spaces

The demo accepts JPEG and PNG uploads.

If you prefer running the demo by yourself, you will need an extra dependency (Streamlit) for the app to run:

pip install -e ".[demo]"

You can then easily run your app in your default browser by running:

streamlit run demo/app.py

torchcam_demo

Visualization script

An example script is provided for you to benchmark the heatmaps produced by multiple CAM approaches on the same image:

python scripts/cam_example.py --arch resnet18 --class-idx 232 --rows 2

gradcam_sample

All script arguments can be checked using python scripts/cam_example.py --help

Performance benchmarks

The purpose of CAM methods is to provide interpretability and they do so by pointing the biggest influence factors on the model outputs. Ideally the CAM should pinpoint all the visual cues that have any influence of the output classification score. For this, we use two metrics:

  • Increase in Confidence (higher is better): the fraction of inputs for which masking with the CAM increases the probability of the original predicted class.
  • Average Drop (lower is better): the mean relative decrease in that class probability after masking, with increases counted as zero drop.

The table below records historical October 2025 results, obtained with earlier metric and CAM implementations. These values have not been revalidated with the current implementation. The ResNet-18 LayerCAM row has been corrected to match the recorded CSV and original benchmark notebook.

CAM method Arch Average drop (↓) Increase in confidence (↑)
GradCAM resnet18 0.2686 0.2250
GradCAMpp resnet18 0.5271 0.1962
SmoothGradCAMpp resnet18 0.2088 0.2499
LayerCAM resnet18 0.1805 0.2894
GradCAM mobilenet_v3_large 0.2678 0.3483
GradCAMpp mobilenet_v3_large 0.3182 0.2535
SmoothGradCAMpp mobilenet_v3_large 0.2681 0.2678
LayerCAM mobilenet_v3_large 0.2526 0.2882

The recorded protocol used the validation set of imagenette2-320, Resize(256) followed by CenterCrop(224), and target layers layer4 for ResNet-18 and features for MobileNet V3 Large. It evaluated the original predicted class. Masking multiplies the ImageNet-normalized input by the CAM: a zero mask therefore corresponds to the ImageNet mean RGB color, not black. No original random seed was recorded; stochastic CAMs and changes to CAM or metric numerics require a fresh benchmark run.

You can run a new benchmark on your hardware with an explicit seed and weight version as follows:

uv run --extra scripts python scripts/eval_perf.py ~/Downloads/imagenette2-320 LayerCAM --arch mobilenet_v3_large --seed 0 --weights IMAGENET1K_V2

The optional --deletion-insertion evaluation uses a normalized zero baseline for both curves and 20 perturbation steps by default. It generates CAMs separately from the classification metrics, consuming extra random draws that can change subsequent stochastic classification CAMs even with the same seed; compare stochastic methods using the same flags. This differs from the original RISE protocol, which uses a blurred image as its insertion baseline; scores from different protocols are not directly comparable.

All script arguments can be checked using python scripts/eval_perf.py --help

Latency benchmark

You crave for beautiful activation maps, but you don't know whether it fits your needs in terms of latency?

The table below preserves historical October 2021 CPU latency measurements (initial forward pass not included), from the original benchmark commit. They have not been revalidated with current implementations. The GPU column has been retired because the original timer did not synchronize CUDA operations; the current scripts/eval_latency.py synchronizes CUDA before and after timing.

CAM method Arch CPU mean (std)
CAM resnet18 0.14ms (0.03ms)
GradCAM resnet18 40.66ms (1.82ms)
GradCAMpp resnet18 41.61ms (3.24ms)
SmoothGradCAMpp resnet18 239.27ms (7.85ms)
ScoreCAM resnet18 6796.89ms (415.14ms)
XGradCAM resnet18 40.63ms (2.03ms)
LayerCAM resnet18 40.91ms (1.79ms)
CAM mobilenet_v3_large N/A*
GradCAM mobilenet_v3_large 26.64ms (3.46ms)
GradCAMpp mobilenet_v3_large 25.50ms (3.10ms)
SmoothGradCAMpp mobilenet_v3_large 156.25ms (4.89ms)
ScoreCAM mobilenet_v3_large 679.16ms (55.04ms)
XGradCAM mobilenet_v3_large 24.21ms (2.94ms)
LayerCAM mobilenet_v3_large 25.14ms (3.17ms)

*The base CAM method cannot work with architectures that have multiple fully-connected layers

These CPU measurements used 100 iterations on (224, 224) inputs and a laptop with an Intel(R) Core(TM) i7-10750H.

You can run this latency benchmark for any CAM method on your hardware as follows:

uv run --extra scripts python scripts/eval_latency.py SmoothGradCAMpp --device cpu --weights none --output latency.json

Each command runs five fresh processes with one CPU thread. It reports the first call after extractor setup, then runs 10 full CAM warm-up calls before collecting 100 samples. Use --repeat, --threads, --warmup, and --it to change these settings. --weights none uses an untrained model and avoids downloads; omit it to use pretrained weights.

The default --scope extractor excludes the initial model forward pass. --scope end-to-end includes the forward pass and CAM extraction; both exclude image loading, preprocessing, device transfers, and extractor setup. The JSON report saves raw samples, median/p95 latency, resolved layers and weights, versions, revision, and script hash. Summary latency statistics are medians of trial statistics; memory uses the highest peak. Peak RSS is the highest RAM use of a worker through setup and measurement, including output checks. CUDA allocated and reserved memory are reported separately. RSS is unavailable on Windows.

ViT and Swin spatial methods preserve channel-first targets and reshape token/channel-last targets. In eval_latency.py, LeGrad needs an explicit transformer block (--target-layer encoder.layers.encoder_layer_11 for vit_b_16), and RefineCAM needs at least two repeated --target-layer arguments. cam_example.py uses --target for both methods: --target encoder.layers.encoder_layer_11 for LeGrad or --target layer3,layer4 for RefineCAM. Targets in one extractor must share a tensor layout.

All script arguments can be checked using python scripts/eval_latency.py --help

Example notebooks

Explore these runnable notebooks, hosted in frgfm/notebooks:

Each notebook includes setup cells that pin a development snapshot with fixes after TorchCAM 0.5.0. Run those cells to use the intended implementation.

Notebook Use case Run
Quicktour Extract CAMs, create overlays, and fuse multiple layers Colab
Debug a prediction Compare predicted and expected classes and save an evidence bundle Colab
Vision Transformers Explain a torchvision ViT with LeGrad and token-based GradCAM Colab
Latency benchmark Measure CAM extraction and end-to-end latency on your hardware Colab
Performance benchmark Evaluate confidence and deletion/insertion faithfulness with an explicit baseline Colab

Citation

If you wish to cite this project, feel free to use this BibTeX reference:

@misc{torcham2020,
    title={TorchCAM: class activation explorer},
    author={François-Guillaume Fernandez},
    year={2020},
    month={March},
    publisher = {GitHub},
    howpublished = {\url{https://github.com/frgfm/torch-cam}}
}

Contributing

Feeling like extending the range of possibilities of CAM? Or perhaps submitting a paper implementation? Any sort of contribution is greatly appreciated!

You can find a short guide in CONTRIBUTING to help grow this project!

License

Distributed under the Apache 2.0 License. See LICENSE for more information.

FOSSA Status

About

Class activation maps for your PyTorch models (CAM, Grad-CAM, Grad-CAM++, Smooth Grad-CAM++, Score-CAM, SS-CAM, IS-CAM, XGrad-CAM, Layer-CAM, Finer-CAM, LeGrad, RefineCAM)

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2.3k stars

Watchers

10 watching

Forks

Releases

Sponsor this project

Used by

Contributors

Languages