A D2Q9 BGK Lattice Boltzmann Method solver for two-dimensional Couette flow, implemented in both CUDA C++ and serial C++ for performance comparison.
The project solves the classical Couette flow problem: the bottom wall is stationary, the top wall moves with a constant velocity, and the fluid between the walls develops an approximately linear velocity profile.
- D2Q9 lattice Boltzmann formulation
- BGK collision model
- Periodic boundary condition in the streamwise direction
- Stationary bounce-back bottom wall
- Moving-wall bounce-back top wall
- CUDA GPU implementation
- Serial C++ CPU baseline
- CSV output for velocity field and centerline profile
- Analytical Couette profile comparison
- CPU vs CUDA timing benchmark
The simulated flow is Couette flow between two parallel plates.
top wall: u_x = U_wall
fluid domain
bottom wall: u_x = 0
For steady Couette flow, the analytical velocity profile is approximately linear:
u_x(y) = U_wall * y / H
where H is the channel height.
cuda-lbm-couette-benchmark/
├── src/
│ ├── lbm_couette_cuda.cu
│ └── lbm_couette_cpu.cpp
├── scripts/
│ ├── plot_couette.py
│ └── plot_benchmark.py
├── results/
│ └── benchmark_summary.csv
├── figures/
├── README.md
└── .gitignore
Use the x64 Native Tools Command Prompt for VS 2022.
Compile CUDA version:
nvcc src\lbm_couette_cuda.cu -o lbm_couette_cuda.exeCompile CPU version:
cl /O2 /EHsc src\lbm_couette_cpu.cpp /Fe:lbm_couette_cpu.exeRun CUDA version with default settings:
lbm_couette_cuda.exeRun CPU version with default settings:
lbm_couette_cpu.exeDefault case:
Grid: 201 x 101
Timesteps: 10000
U_wall: 0.05
omega: 1.0
Optional command-line arguments:
lbm_couette_cuda.exe nx ny timesteps Uwall omega
lbm_couette_cpu.exe nx ny timesteps Uwall omegaExample:
lbm_couette_cuda.exe 201 101 10000 0.05 1.0
lbm_couette_cpu.exe 201 101 10000 0.05 1.0The CUDA solver writes:
results/field_cuda.csv
results/profile_cuda.csv
The CPU solver writes:
results/profile_cpu.csv
The field output contains:
x, y, rho, ux, uy, speed
The profile output contains:
y, ux_numeric, ux_analytical
After running the CUDA solver, create the figures:
python scripts\plot_couette.pyThis generates:
figures/couette_velocity_profile.png
figures/couette_speed_field.png
figures/couette_velocity_vectors.png
To plot the benchmark comparison:
python scripts\plot_benchmark.pyThis generates:
figures/benchmark_solver_time.png
figures/benchmark_speedup.png
The following timings were measured on the development machine using a serial C++ CPU baseline and CUDA implementation. The timed region includes only the solver loop, not CSV writing or plotting.
| Case | Grid | Timesteps | CPU time | CUDA time | Speedup |
|---|---|---|---|---|---|
| Small | 81 × 41 | 5,000 | 0.795 s | 0.507 s | 1.57× |
| Larger | 201 × 101 | 10,000 | 9.423 s | 1.623 s | 5.81× |
For the larger benchmark, the CUDA solver was approximately 5.8 times faster than the optimized serial CPU version.
The velocity profile from the CUDA solver is compared against the analytical Couette profile. The numerical profile follows the expected approximately linear trend from the stationary bottom wall to the moving top wall.
The larger CUDA test produced:
rho range: 1.00004 to 1.00004
max speed: 0.049653
For U_wall = 0.05, the maximum speed approaching 0.05 confirms that the moving-wall boundary condition is working as expected.
- CUDA performance benefits are problem-size dependent.
- For small grids, CUDA launch and memory overhead reduce the speedup.
- For larger grids, the node-wise parallelism of LBM becomes more beneficial.
- This project uses a simple educational D2Q9 BGK implementation and is intended as a clear GPU-LBM benchmark example.
| Item | Details |
|---|---|
| CPU | Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz |
| GPU | NVIDIA Quadro P2000 |
| GPU memory | 4 GB |
| CUDA Toolkit | 12.8 |
| Compiler | MSVC Build Tools 2022 + NVCC |
| Operating system | Windows |