Run nvidia-smi to confirm the model, memory, driver and power limit of your rented GPU. Check topology and the PCIe link, then run a short burn-in and a bandwidth test. If the device is wrong or a check fails, stop, record the output and contact the provider; do not keep paying for a faulty GPU.
Budget about fifteen minutes for this first check, with the tools already installed; longer diagnostics can follow. Keep the original rental specification beside your terminal, including the variant, GPU count and any promised interconnect.
1. Confirm identity with nvidia-smi
Start with these NVIDIA commands, documented in its SMI reference and nvbandwidth README, accessed 28 September 2026:
nvidia-smi
nvidia-smi -L
nvidia-smi -q
The first gives the overview. NVIDIA documents -L as listing models and UUIDs, and -q as requesting all attributes. Save the UUID with your results so support can identify the device.
NVIDIA's SMI reference says the displayed CUDA version is the driver's maximum supported version, not the installed toolkit. A search for "nvidia smi" leads to the same tool, but the executable needs the hyphen.
For repeatable records, NVIDIA's query guide, dated 29 September 2021, documents this CSV logging command. Stop it with Ctrl+C after collecting your samples.
nvidia-smi --query-gpu=timestamp,name,pci.bus_id,driver_version,pstate,pcie.link.gen.max,pcie.link.gen.current,temperature.gpu,utilization.gpu,utilization.memory,memory.total,memory.free,memory.used --format=csv -l 5
NVIDIA also documents nvidia-smi --help-query-gpu for field help. Use the following reading key for the default display. It follows NVIDIA's SMI, NVML and query documentation, plus the nvitop and Microsoft Virtual Client references, accessed 28 September 2026 except the dated query guide.
| Field | Meaning | What to check |
|---|---|---|
| Name / Driver Version | Product and driver identifiers | Match the ordered model; preserve the driver version |
| Persistence-M | Driver state persists after the last client disconnects on Linux | A mode indicator |
| Bus-Id | PCI address | Match the device across reports |
| Disp.A | An initialized display, even without a monitor | Not a workload metric |
| Volatile Uncorr. ECC | Uncorrected errors since driver load | Record before and after testing |
| Fan | Intended fan speed, not measured rotation | Missing data can mean enclosure cooling |
| Temp | GPU core temperature | Watch during load |
| Perf | P0 highest to P12 lowest performance state | Compare idle and loaded states |
| Pwr:Usage/Cap | Power draw and limit in watts | Separate consumption from the cap |
| Memory-Usage | Used and total memory in MiB | Check capacity and existing allocations |
| GPU-Util | Sampling time with at least one kernel executing | Activity, not a throughput score |
| Compute M. | Default: multiple contexts; Exclusive Process: one; Prohibited: none | Check whether computation is permitted |
| MIG M. / N/A | MIG mode / unavailable information | Investigate the allocation type |
NVIDIA's utilization definition makes a busy GPU insufficient evidence of good performance. Continue through the transfer and error checks even if utilization looks high.
Match the exact variant and allocation
The site's GPU specification reference supplies these family values. Treat the component as a family reference, not an acceptance threshold for every variant.
| Spec | H100 | H200 | A100 | L40S | RTX 4090 | B200 |
|---|---|---|---|---|---|---|
| VRAM | 80 to 94 GB | 141 GB | 40 to 80 GB | 48 GB | 24 GB | 180 to 192 GB |
| Memory type | HBM3 | HBM3e | HBM2e | GDDR6 | GDDR6X | HBM3e |
| Memory bandwidth | 3,350 GB/s | 4,800 GB/s | 2,039 GB/s | 864 GB/s | 1,008 GB/s | 8,000 GB/s |
| TDP | 700 W | 700 W | 400 W | 350 W | 450 W | 1,000 W |
| Interconnect | NVLink, PCIe 5.0, InfiniBand | NVLink, PCIe 5.0, InfiniBand | NVLink, PCIe 4.0, InfiniBand | PCIe 4.0 | PCIe 4.0 | NVLink, PCIe 6.0, InfiniBand |
Microsoft's Virtual Client documentation reports memory in MiB. NVIDIA's NVML and SMI references say ECC storage on affected products and driver reservations can reduce available memory. NVIDIA also distinguishes default, configured and enforced power limits. Compare units and the exact variant before treating a difference from the table as a fault.
For example, the site's A100 records include both 40 GB and 80 GB models. Its H100 records distinguish PCIe and SXM variants. Compare the delivered variant with the exact order and any substitutions the provider allows. The NVLink, PCIe and SXM guide explains the listing terminology.
NVIDIA's MIG guide, accessed 28 September 2026, shows indented MIG devices with separate UUIDs under the parent GPU in -L output. It also documents this profile query:
nvidia-smi mig -lgip
If you ordered a whole GPU, resolve an unexpected slice before proceeding. Read MIG and fractional GPUs when comparing the allocation. NVIDIA notes that Ampere MIG utilization can display N/A while CUDA runs, so that field alone cannot establish inactivity.
2. Inspect power, clocks and PCIe under load
Lambda documents the memory and power query below; xCAT documents the permissible power-limit bounds. Both references were accessed 28 September 2026.
nvidia-smi -q -d MEMORY,POWER
nvidia-smi -i 0 --query-gpu=power.min_limit,power.max_limit --format=csv
For GPU 0, compare the configured and enforced limits with the device default in the detailed report. An unexpectedly low cap is a question for the provider, not proof of deliberate throttling. Do not substitute a family TDP for the promised variant's configuration.
Red Hat's parser documentation, accessed 28 September 2026, gives this clock-event query:
/usr/bin/nvidia-smi --query-gpu=name,clocks_event_reasons.active --format=csv,noheader
NVIDIA's NVML event reference, accessed the same day, explains the distinctions: software power capping constrains clocks to the configured limit; software thermal slowdown can reflect GPU or memory temperature limits; idle events reflect downclocking without work. Record active reasons during the burn-in, not just before it.
For PCIe width, Microsoft's Virtual Client reference documents pcie.link.width.current. Use that field with NVIDIA's documented CSV query syntax:
nvidia-smi --query-gpu=pcie.link.width.current --format=csv
Read it alongside the generation fields in the earlier logger. NVIDIA's 2021 query guide says current generation can fall while idle and maximum generation can be host-limited. Investigate a link that stays below the contracted configuration during active transfers. That alone does not prove the GPU itself is defective.
3. Check topology and NVLink status
NVIDIA's SMI reference, accessed 28 September 2026, documents both commands and the topology legend:
nvidia-smi topo -m
nvidia-smi nvlink --status
| Code | NVIDIA's meaning |
|---|---|
| NV# | Bonded connection comprising # NVLinks |
| PIX | One PCIe switch |
| PXB | Multiple PCIe switches without a host bridge |
| PHB | Through a PCIe host bridge |
| NODE | Between host bridges within one NUMA node |
| SYS | Across NUMA nodes through their SMP interconnect |
Read the matrix for every GPU pair your job will use. Compare the reported link states with the topology promised in the rental. Escalate missing expected NVLink connections before starting a distributed job. Do not demand NVLink on every model: NVIDIA's L40S product page, accessed 28 September 2026, explicitly lists no NVLink support.
4. Inspect memory health before stressing the GPU
NVIDIA documents row-remapper reporting in its SMI reference and page-retirement reporting in its guide dated 9 September 2026. Replace the placeholder with the target GPU identifier.
nvidia-smi -q -d ROW_REMAPPER
nvidia-smi -i <target gpu> -q -d PAGE_RETIREMENT
Also inspect ECC counts in the full -q report. NVIDIA's memory-management guide, dated 9 September 2026, says row remapping replaces legacy page retirement starting with Ampere. Pending remaps require a reset to activate; recorded remap counts need not represent remaps already applied in hardware.
Record counts and pending or failure status before and after the workload. NVIDIA's DCGM software-plugin documentation, accessed 28 September 2026, flags pending or failed remapping and pending or excessive page retirement. Treat those findings as escalation evidence, not something to hide by clearing counters.
NVIDIA's Xid guide, accessed 28 September 2026, places driver error reports in system logs. Use its codes together with the dated memory-management guides:
| Xid | Meaning |
|---|---|
| 48 | Uncorrectable double-bit ECC error |
| 63 | Remap entry recorded, or legacy page retired successfully |
| 64 | Remap-entry insertion or legacy page retirement failed |
| 74 | NVLink connection error; a remote device may be responsible |
| 79 | GPU inaccessible over the bus |
| 94 | Contained error requiring affected application restart |
| 95 | Uncontained error requiring GPU reset before application restart |
Preserve the surrounding log messages. Ask the provider to handle required resets and explain recurring errors before you resume work.
5. Run a short server GPU stress test
NVIDIA's DCGM feature overview documents the discovery command below; its diagnostic command reference documents the starting diagnostic level. Both were accessed 28 September 2026:
dcgmi discovery -l
dcgmi diag -r 1
The software-plugin documentation says level 1 checks deployment without compute stress. It can identify missing device nodes, permissions and device-cgroup restrictions. A pass here is permission to continue checking, not your burn-in result.
| DCGM level | Current command reference's coverage |
|---|---|
| 1 | Software deployment |
| 2 | Adds GPU memory and PCIe tests |
| 3 | Adds sustained compute, memory bandwidth, targeted stress and power, plus available conditional tests |
| 4 | Adds memtest and pulse_test |
For a supported data-center GPU, continue with NVIDIA's documented dcgmi diag -r 2. Choose longer diagnostics separately from the initial acceptance window. NVIDIA's current reference only guarantees non-data-center support at level 1 unless more is explicitly documented; do not assume higher levels cover a consumer GPU.
The gpu-burn README, accessed 28 September 2026, documents gpu_burn [OPTIONS] [TIME], with -i N selecting GPU N. For a short gpu burn, inspect its help and choose a brief TIME value before invoking that syntax:
gpu_burn -h
gpu_burn -i 0 TIME
Replace TIME before execution. Run it on the allocated GPU with your own workload stopped. Keep NVIDIA's documented monitor open in another terminal:
nvidia-smi dmon
Repeat the health queries afterward. For interactive observation, the nvtop README describes charts for utilization, temperature, power, clocks and PCIe traffic. The gpustat README documents gpustat --watch for a compact refreshing display. Both were accessed 28 September 2026.
6. Test transfers, then NCCL for multiple GPUs
NVIDIA's nvbandwidth README, accessed 28 September 2026, documents these host-transfer and device-transfer tests. Use the device-to-device test when your allocation includes multiple GPUs.
nvbandwidth -t host_to_device_memcpy_ce device_to_host_memcpy_ce
./nvbandwidth -t device_to_device_memcpy_write_ce
Request a provider baseline for the same test and configuration if results concern you. There is no universal pass number here. Keep the command with the results so support can reproduce the comparison.
NVIDIA's nccl-tests README, accessed 28 September 2026, documents this single-node example for eight GPUs. Set -g to your allocated GPU count before running your NCCL test.
./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8
The README says NCCL tests check collective correctness and performance, and require an MPI-enabled build for multiple processes and nodes. For multi-node acceptance, request the provider's MPI launch command and run across the allocated nodes. A single-node result leaves that check unfinished. Use the InfiniBand and RoCE guide to frame the network questions.
NVIDIA's NCCL performance guide, accessed 28 September 2026, defines algbw as data size divided by elapsed time. It normalizes busbw for the collective; AllReduce uses algbw * (2*(n-1)/n) for n ranks. The guide distinguishes small-message overhead from large-message bandwidth. Keep message sizes and rank count with your report.
7. Resolve visibility failures or reject the instance
If the container cannot see the expected GPUs, NVIDIA's Container Toolkit documentation, accessed 28 September 2026, identifies NVIDIA_VISIBLE_DEVICES and Docker's --gpus as exposure controls. It requires the utility capability for SMI and compute for CUDA. NVIDIA's CUDA environment reference says CUDA_VISIBLE_DEVICES separately controls application visibility and enumeration order. Check those settings before declaring missing hardware.
For a hardware or configuration failure, stop and preserve the model, UUID, driver, command, timestamps and error output. NVIDIA's Xid guide documents this evidence collector:
sudo nvidia-bug-report.sh
Send the resulting nvidia-bug-report.log.gz to provider support with the original order. Request correction or replacement. Follow the provider's controls to end billing for a faulty instance; stopping your test is not a billing action.
Use the GPU server rental checklist for the next order and the H100 rental page when comparing replacements. Accept the server only after identity, health and the links your workload needs check out.
Sources
- NVIDIA SMI reference, accessed 28 September 2026.
- NVIDIA useful nvidia-smi queries, 29 September 2021.
- NVIDIA NVML device queries, accessed 28 September 2026.
- nvitop documentation, accessed 28 September 2026.
- Microsoft Virtual Client NVIDIA monitor, accessed 28 September 2026.
- NVIDIA MIG guide, accessed 28 September 2026.
- Lambda GPU troubleshooting, accessed 28 September 2026.
- xCAT GPU management, accessed 28 September 2026.
- Red Hat Insights NVIDIA parser, accessed 28 September 2026.
- NVIDIA clock-event reasons, accessed 28 September 2026.
- NVIDIA L40S, accessed 28 September 2026.
- NVIDIA page-retirement visibility, 9 September 2026.
- NVIDIA row remapping, 9 September 2026.
- NVIDIA memory-health statistics, 9 September 2026.
- NVIDIA Xid errors, accessed 28 September 2026.
- NVIDIA DCGM discovery, accessed 28 September 2026.
- NVIDIA DCGM diagnostic command reference, accessed 28 September 2026.
- NVIDIA DCGM software checks, accessed 28 September 2026.
- gpu-burn README, accessed 28 September 2026.
- NVTOP README, accessed 28 September 2026.
- gpustat README, accessed 28 September 2026.
- NVIDIA nvbandwidth, accessed 28 September 2026.
- NVIDIA NCCL tests, accessed 28 September 2026.
- NVIDIA NCCL performance definitions, accessed 28 September 2026.
- NVIDIA Container Toolkit controls, accessed 28 September 2026.
- NVIDIA CUDA environment variables, accessed 28 September 2026.