Chisel – Profile GPU Kernels Without a GPU (Nvidia and AMD)
Chisel – Profile GPU Kernels Without a GPU (Nvidia and AMD)
We built Chisel to make GPU kernel profiling hardware-free. It lets you run chisel profile kernel.cu and get full Nsight/Ncompute or rocprofv3 reports without a GPU needed. It spins up remote H100, L40S, or MI300X machines (via DigitalOcean for now, but gonna expand backends soon), runs your code, and gives you back detailed traces (kernel timings, memory transfers, API calls, etc). Everything is CLI-based and designed for iterative dev—profiling takes \~1–2 minutes per run. For example: # Profile a PyTorch training script on H100 with Nsight Systems chisel profile --nsys train.py # Profile a HIP kernel on MI300X with system trace chisel profile --rocprofv3="--sys-trace" matrix_add.cpp Repo: https://github.com/Herdora/chisel PyPI: pip install chisel-cli Would love feedback! especially from anyone building custom kernels, ML layers, or low-level GPU ops.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
trex_exporter – Prometheus Exporter for T-Rex Nvidia GPU Miner
Gource visualizations rendered without a GPU
Lambda Echelon GPU Cluster
GPU PaaS
Slurmq – GPU quota enforcement for Slurm
GPU Accelerated PDAL
Tilery-VM – run Nvidia cuTile GPU kernels on a CPU, no GPU required
Realtime LLM Chat on an 8GB Nvidia GPU
Profile GPU Kernels with One Command, Zero GPU Setup
Nvcachetools – See compiled shader code for Nvidia GPU, using the cache