How to properly apply thread synchronization in CUDA app?
如何正确应用在CUDA应用程序线程同步吗?
Is this a CUDA thread synchronization issue or something else?
这是CUDA线程同步问题还是其他什么?
Appendix B lists the mathematical functions supported in CUDA.
附录b列举cuda中支持的数学函数。
How to query the current performance state of your GPU with CUDA?
如何查询你的GPU使用CUDA的当前性能状态?
You likely created a new CPP file using "CUDA C Bitreverse Application" template.
你可能会创建一个新的CPP文件使用CUDAC倒位应用模板。
In this paper, we implement an efficient matrix multiplication on GPU using NVIDIA's CUDA.
本文使用NVIDIA的CUDA在GPU上实现了一个高效的矩阵乘法。
Using CUDA C language, using CUDA texture memory, image stretching parallel implementation.
说明:使用CUDAC语言,利用CUDA纹理内存,实现图像拉伸的并行实现。
We do take every opportunity to discuss the ability to run CUDA with anyone who's interested.
但我们的确在抓紧每个机会与那些对CUDA感兴趣的人讨论运行CUDA的能力问题。
This document is divided into the following chapters: chapter 1 is an introduction to CUDA and GPU.
本文档分为以下几个章节:第1章是CUDA和GPU的简介。
That being said, as of CUDA 4.0 by default there is one context created per process and not per thread.
也就是说,默认4.0CUDA技术的每个过程,而不是有一个上下文创建每个线程。
Achieve a highly paralleled algorithm to calculate the simplification error of triangular meshes by using CUDA.
利用CUDA实现了高度并行化的网格模型简化误差计算算法。
The result showed that CUDA could speed up calculation and be well used in real-time target tracking on upper computer.
结果表明,CUDA的应用使上位机目标跟踪的实时性得到了很大提升,可以将其应用于其它众多领域。
Multiple NPN240s can be linked to single or multiple hosts to create multi-node CUDA GPU clusters capable of thousands of GFLOPS.
多个NPN240处理器可以链接到一个或多个主机,建立多节点CUDAGPU集群,峰值可达数千gflops。
After experiments, comparing CPU 's computing power can be found, CUDA' s ability to process data in parallel is very strong.
在经过实验之后,对比CPU的计算能力可以发现,CUDA在并行处理数据的能力非常强大。
Abstract CUDA is a parallel computing architecture introduced by NVIDIA, it mainly used for large scale data-intensive computing.
摘要CUDA是一种由NVIDIA推出的并行计算架构,非常适合大规模数据密集型计算。
The CUDA driver and Toolkit installation are required before running the precompiled examples or compiling the example source code.
必需安装CUDA驱动和CUDA工具包,此后才可运行预编译的例程或编译样例源代码。
Each CUDA context has it's own virtual memory space, therefore you can not use a pointer from one context inside an another context.
每个CUDA上下文都有它自己的虚拟内存空间,因此你不能使用一个指针从一个上下文在另一个上下文。
CUDA just take full advantage of parallel capability of GPU, which is a kind of scalable parallel computing model launched by the NVIDIA.
CUDA正是为了充分利用GPU的并行功能,由NVIDIA公司推出的可伸缩并行计算模型。
The CUDA application completely runs on the target machine, so the console or UI for the application will be seen on the target machine only.
CUDA应用程序完全运行在目标机器上,所以控制台或用户界面的应用程序将被视为对目标机。
Please note that the CUDA Debugger for Linux has been tested only on 32-bit Red hat Enterprise Linux (RHEL) 5.x but may work on other distros as well.
注意:Linux平台下的CUDA调试程序仅在32位的Linux红帽企业版5 .x (RHEL)上测试通过,可能也支持Linux其他已发行版本。
CUDA gives full play to the advantages of GPU Streaming Multiprocessors Array and greatly improves the efficiency of the parallel computation programs.
倍。CUDA使GPU流处理器阵列的性能得到充分发挥,极大地提高了并行计算程序的效率。
The core part of ray tracing computation will be modified to adapt the advantages and limitations of CUDA so as to amplify the power of parallelization.
光线追踪的核心计算部分则根据CUDA优势与限制进行适应性改造,发挥尽可能大的并行能力。
Any GPU device has a device driver, so targeting it makes more sense than generating CUDA or OpenCL code which would require from users to install other SDKs.
所有GPU设备都有设备驱动,因此针对它来编程更合理,这样会比生成CUDA或者OpenGL的代码更好,因为那还需要用户安装其它的SDK。
Therefore, in order to improve solving efficiency of packing problem fundamentally, we design parallel algorithm with the structure of CUDA based on GPU.
因此,为了从根本上提高布局问题的求解效率,本文采用基于GPU结构的CUDA技术设计并行算法。
The execute model of single instruction, multiple threads (SIMT) of CUDA is very suitable for parallel to execute the same operations for large-scale data;
CUDA的单指令、多线程(SIMT)的执行模型,很适合大型数据上并行执行相同的操作;
We hope you'll take a look at the new CUDA Toolkit 3.0, and learn more about the tools and resources we've got for all NVIDIA developers in the developer Zone.
我们希望您在新的CUDA技术工具包3.0看看,了解的工具,我们已经在开发区所有NVIDIA开发了更多资源。
Proper data structure is designed to solve the problems such as CUDA does not support pointer, dynamically memory allocating and to avoid synchronization as much as possible.
同时设计了相应的数据结构,克服了CUDA没有指针、不能动态申请资源、尽量避免同步操作等问题。
Aiming at the low rate of route planning due to huge, complex 3d data, this paper proposes a 3d data field route planning based on Compute Unified Device Architecture (CUDA).
针对数据量庞大、复杂的三维数据场环境下航路规划速度偏低的问题,提出一种基于统一计算设备架构(CUDA)的三维数据场航路规划方法。
Each CUDA-capable GPU node includes local DDR3 SDRAM as well as a 16-lane PCI Express? gen2 interface to the system backplane, providing maximum data throughput direct to GPU memory.
每个CUDAGPU节点包括本地的DDR3SDRAM以及一个16通道PCI二代系统的背板接口,直接向GPU内存提供最大的数据流量。
Each CUDA-capable GPU node includes local DDR3 SDRAM as well as a 16-lane PCI Express? gen2 interface to the system backplane, providing maximum data throughput direct to GPU memory.
每个CUDAGPU节点包括本地的DDR3SDRAM以及一个16通道PCI二代系统的背板接口,直接向GPU内存提供最大的数据流量。
应用推荐