Appearance
Debugging
Check where the run happens
Find out whether a device build is running on the host.
fortran
print '(A,L1)', 'host fallback: ', dev_is_host_fallback()text
$ debugging fallback
host fallback: TThe documentation runs on the host on purpose. On a GPU node, T means that the device is not visible (driver, CUDA_VISIBLE_DEVICES, a GPU-less build of the compiler runtime): dev_init also writes a warning on standard error, unless the host was requested.
Find a leaked allocation
Label the allocations and list the live ones at teardown.
fortran
call dev_alloc(fptr_dev=u_dev, ubounds=[100], ierr=ierr, label='u')
if (ierr /= 0) error stop 'device allocation failed'
call dev_alloc(fptr_dev=halo_dev, ubounds=[8], ierr=ierr, label='halo buffer')
if (ierr /= 0) error stop 'device allocation failed'
call dev_free(u_dev)
! ... halo_dev is never freed
call dev_alloc_report() ! at teardown: what is still allocated?text
$ debugging leak
FUNDAL live allocation: address=0x<address> bytes=64 device=0 label="halo buffer"
FUNDAL live allocations: 1 (64 bytes)The address (shown here as a placeholder) changes from run to run; labels are truncated to 32 characters.
Catch a double free
Free through ierr to get an error code instead of a second free.
fortran
alias => a_dev
call dev_free(a_dev, ierr=ierr) ! frees the buffer and nullifies a_dev, but not alias
print '(A,I0)', 'first dev_free: ', ierr
call dev_free(alias, ierr=ierr) ! the same buffer again: detected, nothing is freed
print '(A,I0,A,L1)', 'second dev_free: ', ierr, ', FUNDAL_ERR_NOT_REGISTERED: ', ierr == FUNDAL_ERR_NOT_REGISTEREDtext
$ debugging double-free
first dev_free: 0
second dev_free: 103, FUNDAL_ERR_NOT_REGISTERED: TWARNING
Without ierr, the default policy (warn) writes a warning and then frees the pointer as FUNDAL did before the registry existed: a real double free, which may crash the runtime. Use ierr or the error policy.
A dev_id that contradicts the buffer
dev_free frees a buffer on the device where it was allocated; a dev_id that says otherwise is reported.
fortran
call dev_free(a_dev, dev_id=mydev+1, ierr=ierr) ! the buffer lives on mydev
call dev_get_alloc_stats(allocs=allocs)
print '(A,I0,A,L1,A,I0)', 'dev_free(dev_id=mydev+1): ', ierr, ', FUNDAL_ERR_DEV_ID_MISMATCH: ', &
ierr == FUNDAL_ERR_DEV_ID_MISMATCH, ', live allocations: ', allocs
call dev_free(a_dev, ierr=ierr) ! the buffer is freed on the device where it lives
call dev_get_alloc_stats(allocs=allocs)
print '(A,I0,A,I0)', 'dev_free: ', ierr, ', live allocations: ', allocstext
$ debugging dev-id
dev_free(dev_id=mydev+1): 104, FUNDAL_ERR_DEV_ID_MISMATCH: T, live allocations: 1
dev_free: 0, live allocations: 0Make misuse fatal without changing the code
Set FUNDAL_REGISTRY=error: a misuse of dev_free without ierr stops the program with an error.
fortran
alias => a_dev
call dev_free(a_dev)
print '(A)', 'freeing the buffer a second time, through alias'
call dev_free(alias) ! FUNDAL_REGISTRY=error: error stop
print '(A)', 'not reached'text
$ FUNDAL_REGISTRY=error debugging misuse
freeing the buffer a second time, through alias
[exit status 1]The message (FUNDAL error: dev_free: pointer not allocated by FUNDAL ...) goes to standard error. dev_set_registry_policy in the code takes precedence over the variable.
Debug kernels on a GPU
| Tool | Use |
|---|---|
nvfortran -Minfo=accel | which loops became kernels, and what data they move |
NV_ACC_NOTIFY=3 | trace kernel launches and data transfers (nvfortran) |
NV_ACC_DEBUG=1 | verbose OpenACC runtime output (nvfortran) |
compute-sanitizer ./program | out-of-bounds and illegal accesses on NVIDIA GPUs |
gfortran -fcheck=all -g | bounds checking on the host: runs the kernels on the host, catches indexing errors |
OMP_TARGET_OFFLOAD=MANDATORY | OpenMP: fail instead of running the kernels on the host |