Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -25,12 +25,15 @@ internal fun chooseCostAwarePath(
var bestLength = 1
val match = matches[offset]
val distance = match and MATCH_DISTANCE_MASK
val distanceSymbol = FIXED_DISTANCE_REVERSE_LOOKUP[distance] and 31
val distanceCost = (FIXED_DISTANCE_TREE[distanceSymbol].toInt() and 0xFF) +
(FIXED_DISTANCE_EXTRA_BITS[distanceSymbol].toInt() and 0xFF)
val maximumLength = match ushr MATCH_DISTANCE_BITS
val containedLength = minOf(maximumLength, size - offset)
val searchedLength = minOf(containedLength, COST_AWARE_LENGTH_SEARCH)

for (length in 3..searchedLength) {
val candidateCost = fixedMatchBitCost(length, distance) + costs[offset + length]
val candidateCost = FIXED_LENGTH_TOKEN_COSTS[length] + distanceCost + costs[offset + length]
if (candidateCost <= bestCost) {
bestCost = candidateCost
bestLength = length
Expand All @@ -39,7 +42,7 @@ internal fun chooseCostAwarePath(

if (maximumLength > searchedLength) {
val nextCost = if (offset + maximumLength <= size) costs[offset + maximumLength] else 0
val candidateCost = fixedMatchBitCost(maximumLength, distance) + nextCost
val candidateCost = FIXED_LENGTH_TOKEN_COSTS[maximumLength] + distanceCost + nextCost
if (candidateCost <= bestCost) {
bestCost = candidateCost
bestLength = maximumLength
Expand All @@ -56,10 +59,8 @@ internal fun fixedLiteralBitCost(literal: Int): Int {
}

internal fun fixedMatchBitCost(length: Int, distance: Int): Int {
val lengthSymbol = FIXED_LENGTH_REVERSE_LOOKUP[length] and 31
val distanceSymbol = FIXED_DISTANCE_REVERSE_LOOKUP[distance] and 31
return (FIXED_LENGTH_TREE[257 + lengthSymbol].toInt() and 0xFF) +
(FIXED_LENGTH_EXTRA_BITS[lengthSymbol].toInt() and 0xFF) +
return FIXED_LENGTH_TOKEN_COSTS[length] +
(FIXED_DISTANCE_TREE[distanceSymbol].toInt() and 0xFF) +
(FIXED_DISTANCE_EXTRA_BITS[distanceSymbol].toInt() and 0xFF)
}
Expand All @@ -68,3 +69,9 @@ internal fun fixedMatchBitCost(length: Int, distance: Int): Int {
internal const val COST_AWARE_WINDOW_SIZE = 262_144
// Price every short match length plus the longest available match.
private const val COST_AWARE_LENGTH_SEARCH = 64

private val FIXED_LENGTH_TOKEN_COSTS = IntArray(259) { length ->
val lengthSymbol = FIXED_LENGTH_REVERSE_LOOKUP[length] and 31
(FIXED_LENGTH_TREE[257 + lengthSymbol].toInt() and 0xFF) +
(FIXED_LENGTH_EXTRA_BITS[lengthSymbol].toInt() and 0xFF)
}
79 changes: 79 additions & 0 deletions performance/2026-09-09-costspeed/RESULTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# Reuse fixed token prices in the level 9 parser

Precompute the 259 fixed length-token prices and calculate each match distance price once per parser offset. Candidate order and ties are unchanged. This targets level 9 CPU cost; it does not improve compressed size.

Base: `fe2ff51`, `dev-1.1.0`. Measured on September 9, 2026 on Linux x86_64, Intel Core i7-11800H. JVM uses JBR 17.0.14; Node uses 22.22.1.

Only KFlate was timed. Kompress sizes and timings come unchanged from the trusted September 8 archive in `performance/history.json`. Benchmark setup may produce Kompress streams for decoder validation outside timing.

## Compressed size

The JVM sweep compresses all seven tracked corpora at every listed level. Every output must round-trip through both the JDK RAW inflater and KFlate. Byte counts are deterministic. Single-run durations in the sweep JSON are diagnostic, not throughput benchmarks.

Output reduction is `(beforeBytes - afterBytes) / beforeBytes`. Totals weight each corpus by its compressed bytes; they are not averages of percentage changes.

| Level | Before bytes | After bytes | Output reduction |
| ---: | ---: | ---: | ---: |
| 0 | 76,379,868 | 76,379,868 | 0.0000% |
| 1 | 40,222,941 | 40,222,941 | 0.0000% |
| 2 | 39,269,936 | 39,269,936 | 0.0000% |
| 3 | 38,906,551 | 38,906,551 | 0.0000% |
| 4 | 38,333,169 | 38,333,169 | 0.0000% |
| 5 | 38,293,272 | 38,293,272 | 0.0000% |
| 6 | 37,882,044 | 37,882,044 | 0.0000% |
| 7 | 37,868,330 | 37,868,330 | 0.0000% |
| 8 | 37,737,043 | 37,737,043 | 0.0000% |
| 9 | 37,251,418 | 37,251,418 | 0.0000% |

### Level 6

| Corpus | Before bytes | After bytes | Reduction | Saved JVM Kompress bytes |
| --- | ---: | ---: | ---: | ---: |
| simpleText | 84 | 84 | 0.0000% | 84 |
| text | 506,455 | 506,455 | 0.0000% | 505,318 |
| model3D | 2,153 | 2,153 | 0.0000% | 2,149 |
| Rainier.bmp | 3,283,450 | 3,283,450 | 0.0000% | 3,275,337 |
| Maltese.bmp | 7,158,472 | 7,158,472 | 0.0000% | 7,096,685 |
| Sunrise.bmp | 26,838,825 | 26,838,825 | 0.0000% | 26,698,992 |
| compressed_MVT.pbf | 92,605 | 92,605 | 0.0000% | 91,408 |

### Level 9

| Corpus | Before bytes | After bytes | Reduction | Saved JVM Kompress bytes |
| --- | ---: | ---: | ---: | ---: |
| simpleText | 84 | 84 | 0.0000% | 84 |
| text | 492,523 | 492,523 | 0.0000% | 503,400 |
| model3D | 2,153 | 2,153 | 0.0000% | 2,149 |
| Rainier.bmp | 3,258,570 | 3,258,570 | 0.0000% | 3,270,382 |
| Maltese.bmp | 7,034,092 | 7,034,092 | 0.0000% | 7,090,887 |
| Sunrise.bmp | 26,372,269 | 26,372,269 | 0.0000% | 26,652,810 |
| compressed_MVT.pbf | 91,727 | 91,727 | 0.0000% | 91,400 |

## Compression timing

Development measurements use three one-second warmups and five one-second measurement iterations. JVM has one fresh fork per corpus; Native has one process per corpus; Wasm uses one runner. These shorter runs detect broad tradeoffs and are not release-grade speed claims. Raw samples, errors and confidence intervals are retained. No decompression speed claim is made.

### jvm, level 9

| Corpus | Before ms | After ms | Before/after speedup | Saved Kompress/after speedup |
| --- | ---: | ---: | ---: | ---: |
| simpleText | 0.0272 | 0.0274 | 0.991x | 0.148x |
| text | 192.2328 | 186.2983 | 1.032x | 0.399x |
| model3D | 0.0969 | 0.0946 | 1.024x | 0.317x |
| Rainier.bmp | 2551.1684 | 2307.4431 | 1.106x | 0.048x |
| Maltese.bmp | 1206.9043 | 1168.8631 | 1.033x | 0.339x |
| Sunrise.bmp | 11385.8072 | 11394.6151 | 0.999x | 0.148x |
| compressed_MVT.pbf | 9.3247 | 8.8060 | 1.059x | 0.472x |


## Validation and reproduction

Linux JVM tests passed. All 70 corpus/level outputs round-tripped through JDK and KFlate. SHA-256 hashes matched the baseline for all 70 outputs. All seven level 9 JMH compression cases completed. No Native or Wasm timing claim is made.

Tracked source diff SHA-256: `65031a00c582ebf245a4f16dfd52107b274d0b638cdf047b18f58120a552ebd9`. Added source/test files are present in this commit.

Linux build: `ANDROID_HOME=/home/rafael/Android/Sdk ./gradlew :kflate:jvmTest :kflate:jvmBenchmarkBenchmarkJar --no-parallel --max-workers=1`.

For level 9 timing, set `BENCHMARK_COMPRESSION_LEVEL` to 9 in `BenchmarkState.kt` before building; leave it at 6 otherwise. Use JBR 17.0.14 and the generated JMH jar with the main and benchmark classes on its classpath. Run `org.openjdk.jmh.Main` with the relevant `CompressionBenchmarks` method filter and `-wi 3 -i 5 -w 1s -r 1s -f 1 -foe true -rf json -rff result.json`. Only KFlate benchmark methods should be selected.

Compile `RatioSweep.java` against the KFlate classes and generated benchmark jar, then run it with `kflate/src/jvmTest/resources FIRST_LEVEL LAST_LEVEL`. It checks both JDK and KFlate decompression and records deterministic sizes and output hashes.
39 changes: 39 additions & 0 deletions performance/2026-09-09-costspeed/RatioSweep.java
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
import com.rafambn.kflate.KFlate;
import com.rafambn.kflate.compression.Raw;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Arrays;
import java.security.MessageDigest;
import java.util.HexFormat;
import java.util.zip.Inflater;

public class RatioSweep {
public static void main(String[] args) throws Exception {
String[] names = {"simpleText", "text", "model3D", "Rainier.bmp", "Maltese.bmp", "Sunrise.bmp", "compressed_MVT.pbf"};
int first = args.length > 1 ? Integer.parseInt(args[1]) : 0;
int last = args.length > 2 ? Integer.parseInt(args[2]) : 9;
for (String name : names) {
byte[] input = Files.readAllBytes(Path.of(args[0], name));
for (int level = first; level <= last; level++) {
long start = System.nanoTime();
byte[] output = KFlate.INSTANCE.compress(input, new Raw(level, null));
double ms = (System.nanoTime() - start) / 1e6;
Inflater inflater = new Inflater(true);
try {
inflater.setInput(output);
byte[] decoded = new byte[input.length + 1];
int count = inflater.inflate(decoded);
if (!inflater.finished() || count != input.length || !Arrays.equals(input, Arrays.copyOf(decoded, count))) {
throw new AssertionError("Invalid output: " + name + " level " + level);
}
} finally {
inflater.end();
}
if (!Arrays.equals(input, KFlate.INSTANCE.decompress(output, new com.rafambn.kflate.decompression.Raw(null, null)))) {
throw new AssertionError("KFlate round trip: " + name + " level " + level);
}
System.out.printf(java.util.Locale.ROOT, "{\"corpus\":\"%s\",\"level\":%d,\"originalSizeBytes\":%d,\"compressedSizeBytes\":%d,\"singleRunMs\":%.6f,\"sha256\":\"%s\"}%n", name, level, input.length, output.length, ms, HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(output)));
}
}
}
}
Loading