diff options
| author | Joshua Bakita <bakitajoshua@gmail.com> | 2025-05-05 03:53:01 -0400 |
|---|---|---|
| committer | Joshua Bakita <bakitajoshua@gmail.com> | 2025-05-05 03:53:13 -0400 |
| commit | 293430fcb5d4013b573556c58457ee706e482b7f (patch) | |
| tree | 9328fa680f55b4e1a08d24714275b8437be3be5d /README.md | |
| parent | 494df296bf4abe9b2b484bde1a4fad28c989afec (diff) | |
Snapshot for ECRTS'25 artifact evaluation
Diffstat (limited to 'README.md')
| -rw-r--r-- | README.md | 10 |
1 files changed, 9 insertions, 1 deletions
| @@ -59,6 +59,7 @@ Not all these TPCs will necessarially be enabled in every GPC. | |||
| 59 | Use `cat gpcX_tpc_mask` to get a bit mask of which TPCs are disabled for GPC X. | 59 | Use `cat gpcX_tpc_mask` to get a bit mask of which TPCs are disabled for GPC X. |
| 60 | A set bit indicates a disabled TPC. | 60 | A set bit indicates a disabled TPC. |
| 61 | This API is only available on enabled GPCs. | 61 | This API is only available on enabled GPCs. |
| 62 | Bits greater than the number of on-chip TPCs per GPC should be ignored (it may appear than non-existent TPCs are "disabled"). | ||
| 62 | 63 | ||
| 63 | Example usage: To get the number of on-chip SMs on Volta+ GPUs, multiply the return of `cat num_gpcs` with `cat num_tpc_per_gpc` and multiply by 2 (SMs per TPC). | 64 | Example usage: To get the number of on-chip SMs on Volta+ GPUs, multiply the return of `cat num_gpcs` with `cat num_tpc_per_gpc` and multiply by 2 (SMs per TPC). |
| 64 | 65 | ||
| @@ -83,6 +84,13 @@ Use `echo Z > runlistY/switch_to_tsg` to switch the GPU to run only the specifie | |||
| 83 | 84 | ||
| 84 | Use `echo Y > resubmit_runlist` to resubmit runlist Y (useful to prompt newer GPUs to pick up on re-enabled channels). | 85 | Use `echo Y > resubmit_runlist` to resubmit runlist Y (useful to prompt newer GPUs to pick up on re-enabled channels). |
| 85 | 86 | ||
| 87 | ## Error Interpretation | ||
| 88 | First check the kernel log to see if in includes more information about the error. | ||
| 89 | The following conventions are used for certain error codes: | ||
| 90 | |||
| 91 | - EIO, "Input/Output Error," is returned when an operation fails due to a bad register read. | ||
| 92 | - (Other errors may not have a consistent conventional meaning; see the implementation.) | ||
| 93 | |||
| 86 | ## General Codebase Structure | 94 | ## General Codebase Structure |
| 87 | - `nvdebug.h` defines and describes all GPU data structures. This does not depend on any kernel-internal headers. | 95 | - `nvdebug.h` defines and describes all GPU data structures. This does not depend on any kernel-internal headers. |
| 88 | - `nvdebug_entry.h` contains module startup, device detection, initialization, and module teardown logic. | 96 | - `nvdebug_entry.h` contains module startup, device detection, initialization, and module teardown logic. |
| @@ -94,4 +102,4 @@ Use `echo Y > resubmit_runlist` to resubmit runlist Y (useful to prompt newer GP | |||
| 94 | 102 | ||
| 95 | - The runlist-printing API does not work when runlist management is delegated to the GPU System Processor (GSP) (most Turing+ datacenter GPUs). | 103 | - The runlist-printing API does not work when runlist management is delegated to the GPU System Processor (GSP) (most Turing+ datacenter GPUs). |
| 96 | To workaround, enable the `FALLBACK_TO_PRAMIN` define in `runlist.c`, or reload the `nvidia` kernel module with the `NVreg_EnableGpuFirmware=0` parameter setting. | 104 | To workaround, enable the `FALLBACK_TO_PRAMIN` define in `runlist.c`, or reload the `nvidia` kernel module with the `NVreg_EnableGpuFirmware=0` parameter setting. |
| 97 | (Eg. on A100: end all GPU-using processes, then `sudo rmmod nvidia_uvm nvidia; sudo modprobe nvidia NVreg_EnableGpuFirmware=0`.) | 105 | (Eg. on A100: end all GPU-using processes, then `sudo rmmod nvidia_drm nvidia_modeset nvidia_uvm nvidia; sudo modprobe nvidia NVreg_EnableGpuFirmware=0`.) |
