Friday, October 1, 2010

GPULib 1.4.0 released!

We are pleased to announce that GPULib 1.4.0 is available from the Tech-X website.

For Windows users, we created an installer which installs one of the following pre-built versions of GPULib 1.4.0:
  1. 32- or 64-bit single precision (compute capability 1.0) built with CUDA 3.1
  2. 32- or 64-bit double precision (compute capability 1.3) built with CUDA 3.1
  3. 32- or 64-bit Fermi (compute capability 2.0) built with CUDA 3.1
  4. 32-bit single precision (compute capability 1.0) built with CUDA 2.3
  5. 32-bit double precision (compute capability 1.3) built with CUDA 2.3
(As before, Windows users can opt to build GPULib from source.)

In addition, the following changes have been made:
  • Now builds with CMake for cross-platform compatibility.
  • Now supports CUDA streams, enabling concurrent execution of multiple kernels.
  • Now supports asynchronous data transfer. 
  • Now leverages new features of IDL 8.0 enabling more seamless integration between the two products. 
  • Includes a variety of new algorithms, such as functions for sorting and large histogramming. 
  • gpuinit() now provides additional information, including GPULib version and model of your CUDA-enabled card, and additionally performs a small allocation to ensure the card is available.
  • Demos updated to fail gracefully if you do not have enough memory on your card, or if your card does not have the compute capability to run the demo.
  • Fixed a bug whereby assigning a regular IDL variable to a slice failed.
  • gpushift() fixed for the case where a dimension is shifted by 0.
  • Fixed problems with gpuInterpolate functions.
  • Fixed memory leak in gpufix().
  • Fixed memory leak in gpufft().
  • MATLAB support has been dropped.

Tuesday, November 17, 2009

GPULib 1.2.2 Released

GPULib 1.2.2 is now available from the Tech-X website

We've made extensive changes to the build system, which is now cleaner and more robust. Full release notes follow.

Build system changes 
- Added --with-extra-nvcc-flags=... to configure which allows extra flags to be passed to nvcc.
- If --prefix is not set, make install will fail gracefully, instead of attempting to install in /.
Fixed --with-matlab-dir=... configure option.
- Added IDL and MATLAB configuration info to config.summary to make it easier to troubleshoot problem.
- Added several missing Windows build files.
- Removed several obsolete files and directories. 
- Running 'make clean' will not affect documentation.
- Running 'make install' will properly build code if not already built.
- Install directory is now laid out properly.
- Fixed "No rule to make target `docs/GPULib_UsersGuide.pdf', needed by `all-am'" error.

IDL  changes  
- Fixed bug whereby FFT was only operating on a single row.
- Fixed bug whereby GPUPOW was not found.
- Corrected bwtest example.
- Fixed time reporting for FDTD demo.
- Fixed typo in FDTD demo README.

MATLAB changes 
- Added potential to specify the device number form gpuInit() and accInit().
- Fixed Bug in gpuSet function.
- Fixed Makefile which was incorrect for 32-bit Linux.

2nd edition of Mort Canty's book uses GPULib

From the CRC Press site for Image Analysis, Classification, and Change Detection in Remote Sensing: With Algorithms for ENVI/IDL, Second Edition:

This popular introduction to the processing of remote sensing imagery has been updated to include coverage of the latest versions of the ENVI software environment. This new edition covers support vector machines and other kernel-based methods. Illustrating many programming examples in the array-oriented language ID, the text includes coverage of basic Fourier, wavelet, principal components and minimum noise fraction transformations; convolution filters, topographic modeling, image-to-image registration and ortho-rectification; image fusion; supervised and unsupervised land cover classification with neural networks; hyperspectral analysis; multivariate change detection.

I was excited to hear that GPULib was used in this version of the book. Mort says:

In the text I discuss routines for nonlinear principal component analysis, supervised classification and nonlinear clustering, and explain that they can take advantage of GPULib/CUDA, if installed. (I use your routine GPU_DETECT() to check for GPULib).

Friday, July 24, 2009

GPULib docs from ENVI menu

A recent post by Mort Canty provides a handy program that adds an item to ENVI's help menu that will bring up the GPULib docs.

Tuesday, July 14, 2009

GPULib 1.2 released

GPULib 1.2 is available from the Tech-X website. This release focused on improved MATLAB bindings with a few important bug fixes for the IDL bindings along with a few new kernels. Full release notes follow:

Changes/new features in GPULib version 1.2

General

The main focus of this release is on the improved MATLAB bindings. Some new kernels were added since the release of version 1.0.8.

GPULib kernels

gpuAtan2, gpuFmod, gpuPow

IDL bindings

- Support for the new kernels. For the time being, these funtions only support float and double (so no complex types) and no affine transform arguments.
- Added example bwtest.pro showing the use of page-locked variables for fast CPU/GPU data transfer.
- Added finite-different time-domain example demonstrating the use of views for efficient array sub-selection
- Added spectral angle mapper example.
- Bug fixes for decon_hubble example
- improved documentation

MATLAB bindings

MATLAB GPULib version 1.2 has many major changes from the previous release. READ the README!

First and foremost, there are two distinct and completely separate interfaces to the library. They should NEVER be intermingled.

1.) The accArray class replaces the old gpuArray class from the previous release. This interface requires MATLAB R2008a or higher. This interface hass automatic garbage collection, overloaded operators, and overloaded versions of native MATLAB functions, ...
2.) "gpu" interface class can be used with older versions of MATLAB though it's not clear how far back one can go.

The interface was redesigned for speed. The accArray class is about 2.5X faster than the gpuArray interface for many functions tested. Some of the "gpu"-prefixed functions can be up to 10X faster than the gpuArray interface.

MATLAB GPULib has many new functions including
1.) fft, ifft, fft2, and ifft2
2.) Reduction operations, including sum, cumsum, prod, cumprod,... These functions support 1D vectors and 2D matrices currently.
3.) Single (Complex) and Double (Complex) precision versions of Matrix Multiplication, Transpose and Complex Conjugate Transpose.

The accArray class does not support the subsref.m (i.e. b=A(i)), subsasgn.m (i.e. A(i)=b), or array concatentation functions like A=[B; C; D], These will be supported in future releases.

The "gpu" interface supports subscripting through the gpuSubsref, gpuSubsasgn, and gpuSub2ind functions.

Both interfaces support page-locked host memory allocation via cudaMallocHost. This gives the possibility of much faster memory transfer from CPU memory to GPU memory and back.

Both interfaces include more comprehensive native MATLAB-like documentation.

New examples include: bench, bwtest, fdtd, and fftExample.

Friday, May 8, 2009

GPULib slides from VISualize 2009

Here are the slides from Peter Messmer's GPULib talk at VISualize 2009 a couple weeks ago. Peter and I also stayed around after the IDL presentations to do hands on training for GPULib. The training was well attended; it was good to see the interest in GPU computing with IDL. Most people did not bring a laptop, making it less “hands on” than we had originally intended, but it was good to be able to answer individual questions.

Tuesday, April 21, 2009

NVIDIA Releases OpenCL Driver

OpenCL is officially out -- have a look at the press release, or the OpenCL site.