Skip to content

Repository files navigation

Savitzky-Golay Filter Parallel CUDA

This is a simple C++ parallel implementation about Savitzky-Golay Filter. Our work is based on these papers:

Savitzky-Golay Smoothing Filters in parallel C++ version runs in less than 1 second, to be more precisely in 0.3 ms. We performed experiments on ECG text files. The largest on which we relied is about 400,000 samples.

To compile the code we used the NVCC compiler (Here the documentation) and used libraries such as cuBlas and cuSolver.

Here are some results plotted with MATLAB and execution times:

An image of the signal after using the filter. The blue signal is the original one while the green is the filtered one

signalExample

Instead here we have some run times on a file of about 400.000 samples

Execution_time

Installation and Execution

  • Clone this repository and enter it:
git clone https://github.com/Theangelkk/SGF_Filter.git
cd SGF_Filter
  • Execute it using exectuion.sh. This is an example:
./execution.sh SGFilter_parallel_v03 ECG_signal.txt -5 5 3 0
  • These are the parameters the program needs:
./execution.sh SGFilter_parallel_v03 <txt file> <ML> <MR> <Polynomial Order> <Debug>

The first parameter is the text file from which the signal values are taken. The ML and MR parameters are used to indicate the window size (ML + MR + 1). The following parameter indicates the order of the interpolation polynomial. And finally, the debug parameter tells if you want to print the data structures that are created as you go through the code. There are some constraints to be respected for execution:

  1. ML must be less then MR;
  2. Window size must be less then Polynomial Order;
  3. Window size must be less then input signal size;

Evaluation

Below are some evaluations of signals of different sizes using the following input configuration:

ML = -5
MR = 5
Pol_Ord = 3
Debug = 0
Signal size Time (milliseconds)
34 values 296.26 ms
4.001 values 354.50 ms
46.064 values 345.23 ms
73.114 values 350.27 ms
399.150 values 351.83 ms

The performance has been measured on a machine that has the following characteristics:

  • CPU Intel Core i7 860 2.80 GHz
  • RAM 8 GB DDR3
  • GPU Nvidia Quadro K5000
  • HDD 500 GB SATA

The times shown consider the data transfers from the host to the device and vicersa and the allocation of the various data structures. In addition, performance is affected by the number of users using the machine.

The following shows the time of the main parts of the code. Allocation and transfer times are not considered.

QR Factorization Inverse R Pseudoinverse A Compute M Compute Y Total
34 values 0.11 ms 0.30 ms 0.12 ms 0.03 ms 0.11 ms 0.68 ms
4.001 values 0.11 ms 0.32 ms 0.12 ms 0.03 ms 0.12 ms 0.69 ms
46.064 values 0.11 ms 0.31 ms 0.13 ms 0.05 ms 0.20 ms 0.78 ms
73.114 values 0.11 ms 0.30 ms 0.11 ms 0.07 ms 0.26 ms 0.84 ms
399.150 values 0.11 ms 0.31 ms 0.11 ms 0.28 ms 0.92 ms 1.73 ms

From the table it is easy to notice that the times that change are those for the calculation of the matrix M and the final values Y.

Bibliography

Below are the links that helped us in the development of the code:

Contact

For any questions about code, please contact us.

Github Accounts

Theangelkk, pasqualedetrino or GennaroIannuzzo.

Linkedin Accounts

Theangelkk, pasuqaledetrino or GennaroIannuzzo.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages