Use nvidia-smi to Check a Cloud GPU Before Training

Verify your rented GPU's model, memory, power, PCIe and NVLink, then check errors and run a short server stress test before committing to a long training job.

By Faiz Ahmed•
•12 min read

Run nvidia-smi to confirm the model, memory, driver and power limit of your rented GPU. Check topology and the PCIe link, then run a short burn-in and a bandwidth test. If the device is wrong or a check fails, stop, record the output and contact the provider; do not keep paying for a faulty GPU.

Budget about fifteen minutes for this first check, with the tools already installed; longer diagnostics can follow. Keep the original rental specification beside your terminal, including the variant, GPU count and any promised interconnect.

1. Confirm identity with nvidia-smi

Start with these NVIDIA commands, documented in its SMI reference and nvbandwidth README, accessed 28 September 2026:

nvidia-smi
nvidia-smi -L
nvidia-smi -q

The first gives the overview. NVIDIA documents -L as listing models and UUIDs, and -q as requesting all attributes. Save the UUID with your results so support can identify the device.

NVIDIA's SMI reference says the displayed CUDA version is the driver's maximum supported version, not the installed toolkit. A search for "nvidia smi" leads to the same tool, but the executable needs the hyphen.

For repeatable records, NVIDIA's query guide, dated 29 September 2021, documents this CSV logging command. Stop it with Ctrl+C after collecting your samples.

nvidia-smi --query-gpu=timestamp,name,pci.bus_id,driver_version,pstate,pcie.link.gen.max,pcie.link.gen.current,temperature.gpu,utilization.gpu,utilization.memory,memory.total,memory.free,memory.used --format=csv -l 5

NVIDIA also documents nvidia-smi --help-query-gpu for field help. Use the following reading key for the default display. It follows NVIDIA's SMI, NVML and query documentation, plus the nvitop and Microsoft Virtual Client references, accessed 28 September 2026 except the dated query guide.

FieldMeaningWhat to check
Name / Driver VersionProduct and driver identifiersMatch the ordered model; preserve the driver version
Persistence-MDriver state persists after the last client disconnects on LinuxA mode indicator
Bus-IdPCI addressMatch the device across reports
Disp.AAn initialized display, even without a monitorNot a workload metric
Volatile Uncorr. ECCUncorrected errors since driver loadRecord before and after testing
FanIntended fan speed, not measured rotationMissing data can mean enclosure cooling
TempGPU core temperatureWatch during load
PerfP0 highest to P12 lowest performance stateCompare idle and loaded states
Pwr:Usage/CapPower draw and limit in wattsSeparate consumption from the cap
Memory-UsageUsed and total memory in MiBCheck capacity and existing allocations
GPU-UtilSampling time with at least one kernel executingActivity, not a throughput score
Compute M.Default: multiple contexts; Exclusive Process: one; Prohibited: noneCheck whether computation is permitted
MIG M. / N/AMIG mode / unavailable informationInvestigate the allocation type

NVIDIA's utilization definition makes a busy GPU insufficient evidence of good performance. Continue through the transfer and error checks even if utilization looks high.

Match the exact variant and allocation

The site's GPU specification reference supplies these family values. Treat the component as a family reference, not an acceptance threshold for every variant.

SpecH100H200A100L40SRTX 4090B200
VRAM80 to 94 GB141 GB40 to 80 GB48 GB24 GB180 to 192 GB
Memory typeHBM3HBM3eHBM2eGDDR6GDDR6XHBM3e
Memory bandwidth3,350 GB/s4,800 GB/s2,039 GB/s864 GB/s1,008 GB/s8,000 GB/s
TDP700 W700 W400 W350 W450 W1,000 W
InterconnectNVLink, PCIe 5.0, InfiniBandNVLink, PCIe 5.0, InfiniBandNVLink, PCIe 4.0, InfiniBandPCIe 4.0PCIe 4.0NVLink, PCIe 6.0, InfiniBand
Figures from the vendor datasheets: H100, H200, A100, L40S, RTX 4090, B200, checked 13 Sep 2026. "Not published" means the vendor gives no figure.

Microsoft's Virtual Client documentation reports memory in MiB. NVIDIA's NVML and SMI references say ECC storage on affected products and driver reservations can reduce available memory. NVIDIA also distinguishes default, configured and enforced power limits. Compare units and the exact variant before treating a difference from the table as a fault.

For example, the site's A100 records include both 40 GB and 80 GB models. Its H100 records distinguish PCIe and SXM variants. Compare the delivered variant with the exact order and any substitutions the provider allows. The NVLink, PCIe and SXM guide explains the listing terminology.

NVIDIA's MIG guide, accessed 28 September 2026, shows indented MIG devices with separate UUIDs under the parent GPU in -L output. It also documents this profile query:

nvidia-smi mig -lgip

If you ordered a whole GPU, resolve an unexpected slice before proceeding. Read MIG and fractional GPUs when comparing the allocation. NVIDIA notes that Ampere MIG utilization can display N/A while CUDA runs, so that field alone cannot establish inactivity.

2. Inspect power, clocks and PCIe under load

Lambda documents the memory and power query below; xCAT documents the permissible power-limit bounds. Both references were accessed 28 September 2026.

nvidia-smi -q -d MEMORY,POWER
nvidia-smi -i 0 --query-gpu=power.min_limit,power.max_limit --format=csv

For GPU 0, compare the configured and enforced limits with the device default in the detailed report. An unexpectedly low cap is a question for the provider, not proof of deliberate throttling. Do not substitute a family TDP for the promised variant's configuration.

Red Hat's parser documentation, accessed 28 September 2026, gives this clock-event query:

/usr/bin/nvidia-smi --query-gpu=name,clocks_event_reasons.active --format=csv,noheader

NVIDIA's NVML event reference, accessed the same day, explains the distinctions: software power capping constrains clocks to the configured limit; software thermal slowdown can reflect GPU or memory temperature limits; idle events reflect downclocking without work. Record active reasons during the burn-in, not just before it.

For PCIe width, Microsoft's Virtual Client reference documents pcie.link.width.current. Use that field with NVIDIA's documented CSV query syntax:

nvidia-smi --query-gpu=pcie.link.width.current --format=csv

Read it alongside the generation fields in the earlier logger. NVIDIA's 2021 query guide says current generation can fall while idle and maximum generation can be host-limited. Investigate a link that stays below the contracted configuration during active transfers. That alone does not prove the GPU itself is defective.

NVIDIA's SMI reference, accessed 28 September 2026, documents both commands and the topology legend:

nvidia-smi topo -m
nvidia-smi nvlink --status
CodeNVIDIA's meaning
NV#Bonded connection comprising # NVLinks
PIXOne PCIe switch
PXBMultiple PCIe switches without a host bridge
PHBThrough a PCIe host bridge
NODEBetween host bridges within one NUMA node
SYSAcross NUMA nodes through their SMP interconnect

Read the matrix for every GPU pair your job will use. Compare the reported link states with the topology promised in the rental. Escalate missing expected NVLink connections before starting a distributed job. Do not demand NVLink on every model: NVIDIA's L40S product page, accessed 28 September 2026, explicitly lists no NVLink support.

4. Inspect memory health before stressing the GPU

NVIDIA documents row-remapper reporting in its SMI reference and page-retirement reporting in its guide dated 9 September 2026. Replace the placeholder with the target GPU identifier.

nvidia-smi -q -d ROW_REMAPPER
nvidia-smi -i <target gpu> -q -d PAGE_RETIREMENT

Also inspect ECC counts in the full -q report. NVIDIA's memory-management guide, dated 9 September 2026, says row remapping replaces legacy page retirement starting with Ampere. Pending remaps require a reset to activate; recorded remap counts need not represent remaps already applied in hardware.

Record counts and pending or failure status before and after the workload. NVIDIA's DCGM software-plugin documentation, accessed 28 September 2026, flags pending or failed remapping and pending or excessive page retirement. Treat those findings as escalation evidence, not something to hide by clearing counters.

NVIDIA's Xid guide, accessed 28 September 2026, places driver error reports in system logs. Use its codes together with the dated memory-management guides:

XidMeaning
48Uncorrectable double-bit ECC error
63Remap entry recorded, or legacy page retired successfully
64Remap-entry insertion or legacy page retirement failed
74NVLink connection error; a remote device may be responsible
79GPU inaccessible over the bus
94Contained error requiring affected application restart
95Uncontained error requiring GPU reset before application restart

Preserve the surrounding log messages. Ask the provider to handle required resets and explain recurring errors before you resume work.

5. Run a short server GPU stress test

NVIDIA's DCGM feature overview documents the discovery command below; its diagnostic command reference documents the starting diagnostic level. Both were accessed 28 September 2026:

dcgmi discovery -l
dcgmi diag -r 1

The software-plugin documentation says level 1 checks deployment without compute stress. It can identify missing device nodes, permissions and device-cgroup restrictions. A pass here is permission to continue checking, not your burn-in result.

DCGM levelCurrent command reference's coverage
1Software deployment
2Adds GPU memory and PCIe tests
3Adds sustained compute, memory bandwidth, targeted stress and power, plus available conditional tests
4Adds memtest and pulse_test

For a supported data-center GPU, continue with NVIDIA's documented dcgmi diag -r 2. Choose longer diagnostics separately from the initial acceptance window. NVIDIA's current reference only guarantees non-data-center support at level 1 unless more is explicitly documented; do not assume higher levels cover a consumer GPU.

The gpu-burn README, accessed 28 September 2026, documents gpu_burn [OPTIONS] [TIME], with -i N selecting GPU N. For a short gpu burn, inspect its help and choose a brief TIME value before invoking that syntax:

gpu_burn -h
gpu_burn -i 0 TIME

Replace TIME before execution. Run it on the allocated GPU with your own workload stopped. Keep NVIDIA's documented monitor open in another terminal:

nvidia-smi dmon

Repeat the health queries afterward. For interactive observation, the nvtop README describes charts for utilization, temperature, power, clocks and PCIe traffic. The gpustat README documents gpustat --watch for a compact refreshing display. Both were accessed 28 September 2026.

6. Test transfers, then NCCL for multiple GPUs

NVIDIA's nvbandwidth README, accessed 28 September 2026, documents these host-transfer and device-transfer tests. Use the device-to-device test when your allocation includes multiple GPUs.

nvbandwidth -t host_to_device_memcpy_ce device_to_host_memcpy_ce
./nvbandwidth -t device_to_device_memcpy_write_ce

Request a provider baseline for the same test and configuration if results concern you. There is no universal pass number here. Keep the command with the results so support can reproduce the comparison.

NVIDIA's nccl-tests README, accessed 28 September 2026, documents this single-node example for eight GPUs. Set -g to your allocated GPU count before running your NCCL test.

./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8

The README says NCCL tests check collective correctness and performance, and require an MPI-enabled build for multiple processes and nodes. For multi-node acceptance, request the provider's MPI launch command and run across the allocated nodes. A single-node result leaves that check unfinished. Use the InfiniBand and RoCE guide to frame the network questions.

NVIDIA's NCCL performance guide, accessed 28 September 2026, defines algbw as data size divided by elapsed time. It normalizes busbw for the collective; AllReduce uses algbw * (2*(n-1)/n) for n ranks. The guide distinguishes small-message overhead from large-message bandwidth. Keep message sizes and rank count with your report.

7. Resolve visibility failures or reject the instance

If the container cannot see the expected GPUs, NVIDIA's Container Toolkit documentation, accessed 28 September 2026, identifies NVIDIA_VISIBLE_DEVICES and Docker's --gpus as exposure controls. It requires the utility capability for SMI and compute for CUDA. NVIDIA's CUDA environment reference says CUDA_VISIBLE_DEVICES separately controls application visibility and enumeration order. Check those settings before declaring missing hardware.

For a hardware or configuration failure, stop and preserve the model, UUID, driver, command, timestamps and error output. NVIDIA's Xid guide documents this evidence collector:

sudo nvidia-bug-report.sh

Send the resulting nvidia-bug-report.log.gz to provider support with the original order. Request correction or replacement. Follow the provider's controls to end billing for a faulty instance; stopping your test is not a billing action.

Use the GPU server rental checklist for the next order and the H100 rental page when comparing replacements. Accept the server only after identity, health and the links your workload needs check out.

Sources

Frequently asked questions

Which nvidia-smi command should I run first?▾

Run nvidia-smi for model, memory, driver and power, then nvidia-smi -L for model names and UUIDs. Compare them with the exact variant you ordered.

Does nvidia smi show my installed CUDA toolkit?▾

NVIDIA documents the displayed CUDA version as the maximum supported by the driver. It does not identify the installed toolkit.

Does high GPU utilization prove the server is fast?▾

No. NVIDIA defines utilization as the share of a sampling interval during which a kernel executed; check power, clocks, errors and transfer performance as well.

Which GPU stress test should I use on a server?▾

Start with DCGM deployment checks, then use a supported stress diagnostic or gpu-burn. Follow with nvbandwidth, and use NCCL tests when several GPUs must communicate.

What should I do if a rented GPU fails verification?▾

Stop the workload, save the identity and diagnostic output, and contact the provider. Request repair or replacement before committing to a long run, and end billing for a faulty instance through the provider's controls.

Related Posts