This library introduces AllianceDB , an open-source benchmark suite for evaluating and improving stream operation algorithms on modern hardwares. It includes all codes and experimental traces to reproduce the experiments from our SIGMOD 2021 paper called Parallelizing Intra-Window Join on Multicores: An Experimental Study.
This is a static dump of the code created in 2022 to be uploaded to the ACM website and permanently hosted. Please refer to the canonical DataSysResearch/AllianceDB repository.
run_all.shscript for one-click experiment environment setup and results reproducing.sorting/folder containing scripts and source code for sort-merge join algorithms.hashing/folder containing scripts and source code for hash join algorithms.docs/folder containing useful documentations and figures for understanding both of our framework and datasets.pcm*.cfgconfiguration files to setup event counters used for PCM profiling.cpu-mappings.txtconfiguration file for setting up CPU affinity of threads..
The environment will be automatically configured and all of our experiments can be automatically reproduced by calling the following command with root privileges:
# -d : the experiment results directory
# -c : the L3 cache size of the current CPU
sudo bash run_all.sh -d /data1/xtra -c 19922944-
System specification:
Component Description Processor (w/o HT) Intel(R) Xeon(R) Gold 6126 CPU, 2 (socket) * 12 * 2.6GHz L3 cache size 19MB Memory 64GB, DDR4 2666 MHz OS & Compiler Linux 4.15.0, compile with g++ O3 -
The profiling can be only supported on Intel CPUs, which provide diverse hardware counters.
-
Some PCM profiling results may be incorrect at the first run.
-
You can run any subset of the experiment sections individually by modifying the
exp_sectioninrun_all.sh. -
To run with more than 8 threads, it needs to update the cpu-mapping in
cpu-mapping.txt.
Third-party Libs will be automatically installed in scripts, which contains:
- CMake install, if CMake already installed, make sure the CMake version > 3.10.0.
sudo apt install -y cmake- Tex font rendering:
sudo apt install -y texlive-fonts-recommended texlive-fonts-extra
sudo apt install -y dvipng
sudo apt install -y font-manager
sudo apt install -y cm-super- Python3:
sudo apt install -y python3
sudo apt install -y python3-pip
pip3 install numpy
pip3 install matplotlib- NUMA library
sudo apt install -y libnuma-dev- Zlib
sudo apt install -y zlib1g-dev- python-tk
sudo apt install -y python-tk- perf
sudo apt install -y linux-tools-common
sudo apt install -y linux-tools-`uname -r` # XXX is the kernel version of your linux, use uname -r to check it. e.g. 4.15.0-91-generic
sudo echo -1 > /proc/sys/kernel/perf_event_paranoid # if permission denied, try to run this at root user.
sudo modprobe msr # load msr driverInside the run_all.sh, there are three parameters can be manually configured according to the experiment requirements.
Default parameters are shown as below:
| Parameters | Default | Description |
|---|---|---|
| exp_dir | /data1/xtra | Path to save all results and generate figures |
| L3_CACHE_SIZE | 19922944 (19MB) | Size of L3 cache |
| Experiment Sections | All (e.g. APP_BENCH) | All experiments shown in our paper |
Real world datasets will be downloaded and moved to the exp_dir/datasets automatically by scripts.
We have 4 real datasets that are compressed in datasets.tar.gz. Download and call tar -zvxf datasets.tar.gz to unzip those datasets.
We extracted the useful columns of those datasets, the one is joined key and another is timestamp. The detailed file descriptions of the datasets are shown as below:
| Dataset | Files description |
|---|---|
| DEBS | comments_key32_partitioned.csv : user_id |
| YSB | ad_events.txt : campaign_id |
| Rovio | 1000ms_1t.txt : combined_id | payload|price | timestamp |
| Stock | cj_1000ms_1t.txt : stockid |
All results are in exp_dir/results/.
All figures are in exp_dir/results/figures.