From 293430fcb5d4013b573556c58457ee706e482b7f Mon Sep 17 00:00:00 2001 From: Joshua Bakita Date: Mon, 5 May 2025 03:53:01 -0400 Subject: Snapshot for ECRTS'25 artifact evaluation --- README.md | 10 +++++++++- 1 file changed, 9 insertions(+), 1 deletion(-) (limited to 'README.md') diff --git a/README.md b/README.md index da3e5d7..2889b29 100644 --- a/README.md +++ b/README.md @@ -59,6 +59,7 @@ Not all these TPCs will necessarially be enabled in every GPC. Use `cat gpcX_tpc_mask` to get a bit mask of which TPCs are disabled for GPC X. A set bit indicates a disabled TPC. This API is only available on enabled GPCs. +Bits greater than the number of on-chip TPCs per GPC should be ignored (it may appear than non-existent TPCs are "disabled"). Example usage: To get the number of on-chip SMs on Volta+ GPUs, multiply the return of `cat num_gpcs` with `cat num_tpc_per_gpc` and multiply by 2 (SMs per TPC). @@ -83,6 +84,13 @@ Use `echo Z > runlistY/switch_to_tsg` to switch the GPU to run only the specifie Use `echo Y > resubmit_runlist` to resubmit runlist Y (useful to prompt newer GPUs to pick up on re-enabled channels). +## Error Interpretation +First check the kernel log to see if in includes more information about the error. +The following conventions are used for certain error codes: + +- EIO, "Input/Output Error," is returned when an operation fails due to a bad register read. +- (Other errors may not have a consistent conventional meaning; see the implementation.) + ## General Codebase Structure - `nvdebug.h` defines and describes all GPU data structures. This does not depend on any kernel-internal headers. - `nvdebug_entry.h` contains module startup, device detection, initialization, and module teardown logic. @@ -94,4 +102,4 @@ Use `echo Y > resubmit_runlist` to resubmit runlist Y (useful to prompt newer GP - The runlist-printing API does not work when runlist management is delegated to the GPU System Processor (GSP) (most Turing+ datacenter GPUs). To workaround, enable the `FALLBACK_TO_PRAMIN` define in `runlist.c`, or reload the `nvidia` kernel module with the `NVreg_EnableGpuFirmware=0` parameter setting. - (Eg. on A100: end all GPU-using processes, then `sudo rmmod nvidia_uvm nvidia; sudo modprobe nvidia NVreg_EnableGpuFirmware=0`.) + (Eg. on A100: end all GPU-using processes, then `sudo rmmod nvidia_drm nvidia_modeset nvidia_uvm nvidia; sudo modprobe nvidia NVreg_EnableGpuFirmware=0`.) -- cgit v1.2.2