Problem
There are now two independent traits mapping a Cytnx CPU value type to the CUDA-native type used inside device kernels, with the same content and different spellings:
|
location |
namespace |
specialized on |
to_cuda_t<T> |
src/backend/linalg_internal_gpu/cuTypeCvt.hpp:19-27 |
cytnx::linalg_internal |
std::complex<T> (generic in T) |
internal::gpu_element_type_t<T> |
include/Type.hpp:197-209 (#ifdef UNI_GPU) |
cytnx::internal |
cytnx_complex128, cytnx_complex64 (enumerated) |
Both are the identity for every non-complex type and map complex to cuda::std::complex<...>. Neither references the other.
Current usage: to_cuda_t in 5 files under src/backend/linalg_internal_gpu/; gpu_element_type_t only inside Type.hpp itself, at :318 and :358, to compute cy_typeid_gpu / cy_typeid_gpu_v.
Why it matters
The GPU typed-dispatch migration (#1013, #1003) deliberately keeps the logical dispatch on ordinary Cytnx value types and confines the CUDA-native complex representation to the kernel-launch boundary — that is what to_cuda_t's doc comment says it exists for. Having a second, differently-scoped map for the same job means new dispatch code has to know which one to reach for, and the two can drift: to_cuda already accepts any std::complex<T> while gpu_element_type only accepts the two cytnx complex dtypes.
Related: the trait family is inconsistently CUDA-aware
This surfaced in #1094 review. Within the same family of type traits:
So complex_value_type_t<cuda::std::complex<double>> silently returns cuda::std::complex<double> — the identity branch — rather than double. That is a quiet wrong answer, not a compile error, and it is the kind of thing that makes a GPU output-dtype rule look correct while being wrong. Any code reaching for these traits on GPU scalar types has to know which ones handle cuda::std::complex and which do not.
Suggested scope
- Keep one value-type → CUDA-type map and have the other name alias it (or delete it).
Type.hpp is the natural home since cy_typeid_gpu already depends on it, but either direction is fine as long as there is one definition.
- Decide, and document, whether
complex_value_type should be CUDA-aware like is_complex. If yes, add the cuda::std::complex<T> specialization; if no, say so in the comment so callers do not reach for it on GPU types.
No behavior change intended — this is a consolidation, not a semantics change.
Context
Split out of the #1094 review rather than widening that PR. Noted by @IvanaGyro (reuse of #1092's templates) and @ianmccul (typed dispatch direction).
Problem
There are now two independent traits mapping a Cytnx CPU value type to the CUDA-native type used inside device kernels, with the same content and different spellings:
to_cuda_t<T>src/backend/linalg_internal_gpu/cuTypeCvt.hpp:19-27cytnx::linalg_internalstd::complex<T>(generic inT)internal::gpu_element_type_t<T>include/Type.hpp:197-209(#ifdef UNI_GPU)cytnx::internalcytnx_complex128,cytnx_complex64(enumerated)Both are the identity for every non-complex type and map complex to
cuda::std::complex<...>. Neither references the other.Current usage:
to_cuda_tin 5 files undersrc/backend/linalg_internal_gpu/;gpu_element_type_tonly insideType.hppitself, at:318and:358, to computecy_typeid_gpu/cy_typeid_gpu_v.Why it matters
The GPU typed-dispatch migration (#1013, #1003) deliberately keeps the logical dispatch on ordinary Cytnx value types and confines the CUDA-native complex representation to the kernel-launch boundary — that is what
to_cuda_t's doc comment says it exists for. Having a second, differently-scoped map for the same job means new dispatch code has to know which one to reach for, and the two can drift:to_cudaalready accepts anystd::complex<T>whilegpu_element_typeonly accepts the two cytnx complex dtypes.Related: the trait family is inconsistently CUDA-aware
This surfaced in #1094 review. Within the same family of type traits:
is_complex_impl(Type.hpp:73-81) is specialized for bothstd::complex<T>andcuda::std::complex<T>.complex_value_type(include/utils/complex_arithmetic.hpp:35-45, added in refactor(utils)!: drop namespace-scope complex/builtin operator overloads (#1003) #1092) is specialized only onstd::complex<T>.So
complex_value_type_t<cuda::std::complex<double>>silently returnscuda::std::complex<double>— the identity branch — rather thandouble. That is a quiet wrong answer, not a compile error, and it is the kind of thing that makes a GPU output-dtype rule look correct while being wrong. Any code reaching for these traits on GPU scalar types has to know which ones handlecuda::std::complexand which do not.Suggested scope
Type.hppis the natural home sincecy_typeid_gpualready depends on it, but either direction is fine as long as there is one definition.complex_value_typeshould be CUDA-aware likeis_complex. If yes, add thecuda::std::complex<T>specialization; if no, say so in the comment so callers do not reach for it on GPU types.No behavior change intended — this is a consolidation, not a semantics change.
Context
Split out of the #1094 review rather than widening that PR. Noted by @IvanaGyro (reuse of #1092's templates) and @ianmccul (typed dispatch direction).