Overview

Compress4J is an up-to-date fork of JArchiveLib.

It’s a simple archiving and compression library for Java that provides a thin and easy-to-use API layer on top of the powerful and feature-rich org.apache.commons.compress.

Compress4J is a Java library dedicated to simplifying the handling of various data compressors formats. Its primary goal is to provide developers with a unified and straightforward API for compressing and decompressing files and data streams, regardless of the underlying algorithm being used.

Inspired by the need for a consistent way to interact with different compression techniques, Compress4J abstracts the specific implementations behind common interfaces, making it easier to incorporate compression into Java applications or to switch between compression formats.

Getting Started

To use Compress4J in your Java project, you typically need to include it as a dependency in your build configuration.

  • Maven

  • Gradle Groovy

  • Gradle Kotlin

<dependency>
    <groupId>com.hominux</groupId>
    <artifactId>compress4j</artifactId>
    <version>YOUR_VERSION</version>
</dependency>
dependencies {
    implementation 'com.hominux:compress4j:YOUR_VERSION'
}
dependencies {
    implementation("com.hominux:compress4j:YOUR_VERSION")
}

Java modules

Compress4J is the module com.hominux.compress4j. Commons Compress and Commons IO are implementation details that modular applications do not read; no public signature uses their types. org.tukaani.xz is optional and only needed for the XZ and LZMA formats; modular applications must requires org.tukaani.xz; themselves. com.github.luben.zstd_jni is optional and only needed for the Zstandard formats; modular applications must requires com.github.luben.zstd_jni; themselves. org.brotli:dec is optional and only needed to read Brotli; its jar has no module descriptor, so modular applications start the JVM with --add-modules dec.

Core Concepts

The library’s design centers around a few key components:

  • Compression: a codec and its options, created by a factory such as Compression.gzip().level(9). Every codec is one value type; there is no class per codec.

  • Compressor and Decompressor: final classes that compress to, and decompress from, a file, channel or stream with any Compression. Decompressor detects the codec when you give none and enforces the default limits.

  • ArchiveExtractor: reads an archive as a stream of `ArchiveItem`s (entry plus content) or extracts it safely to a directory. See Streaming entries.

  • ArchiveCreator: writes `EntrySource`s: files, directories and symlinks from disk, memory or another archive.

Each archive format has its own creator and extractor (TarArchiveCreator, ZipArchiveExtractor). A compressed tar is a TarArchiveCreator with a Compression; see tar.

Quick start

try (TarArchiveCreator creator = TarArchiveCreator.builder(Path.of("example.tar.gz"))
        .compression(Compression.gzip())
        .build()) {
    creator.addDirectoryRecursively(Path.of("exampleDir"));
    creator.addFile(Path.of("path/to/file.txt"));
}

try (TarArchiveExtractor extractor =
        TarArchiveExtractor.builder(Path.of("example.tar.gz")).build()) {
    extractor.extract(OUTPUT_DIR);
}

try (Compressor compressor = Compressor.builder(
                Path.of("file.txt.zst"), Compression.zstd().level(6))
        .build()) {
    compressor.write(Path.of("file.txt"));
}

try (Decompressor decompressor =
        Decompressor.builder(Path.of("file.txt.zst")).build()) {
    decompressor.write(Path.of("copy.txt"));
}

Supported Formats

Compress4J reads and writes these archive and compression formats:

  • AR (.a, .ar)

  • ARJ (.arj, read-only)

  • CPIO (.cpio)

  • UNIX dump (.dump, read-only)

  • Deflate (.deflate)

  • Deflate64 (read-only)

  • Brotli (read-only)

  • Bzip2 (.bz2)

  • Gzip (.gz, gzip)

  • Pack200 (.pack, never detected)

  • Snappy (.sz framed, and raw)

  • Tar (.tar, tar.bz2, tbz2, tar.gz, tgz, tar.xz, txz, tar.lzma, tar.lz4, tar.zst, tar.Z read-only)

  • LZ4 (.lz4, framed and block)

  • LZMA (.lzma, never detected)

  • Unix compress (.Z, tar.Z, read-only)

  • XZ (.xz)

  • Zip (.zip)

  • Zstandard (.zst)

Formats marked (read-only) can be extracted or decompressed, as applicable, but not created; see Read-only formats.

Supported codecs

A test pins the Write and Detected columns to the codec catalog the contract tests run against, and the Compression value column to the Compression factories. Detected means Decompressor and the tar extractor pick the codec from the stream’s leading bytes; select the others with their Compression value. Input with no detectable signature passes through Decompressor unchanged.

Codec Read Write Detected Compression value

gzip

✓

✓

✓

Compression.gzip()

bzip2

✓

✓

✓

Compression.bzip2()

xz

✓

✓

✓

Compression.xz()

lzma

✓

✓

Compression.lzma()

lz4-block

✓

✓

Compression.lz4Block()

lz4-framed

✓

✓

✓

Compression.lz4Framed()

zstd

✓

✓

✓

Compression.zstd()

deflate

✓

✓

Compression.deflate()

snappy-framed

✓

✓

✓

Compression.snappyFramed()

snappy-raw

✓

✓

Compression.snappyRaw()

pack200

✓

✓

Compression.pack200()

deflate64

✓

Compression.deflate64()

brotli

✓

Compression.brotli()

z

✓

✓

Compression.unixZ()

Some compression formats may require additional support libraries to be included in your project. Please refer to the Additional Formats for details on any extra dependencies needed for specific formats.

Extracting Untrusted Archives

Path traversal is always rejected: entry paths are resolved canonically against the output directory, so both ../ entries and writes that would follow a symlink out of the output directory fail with UnsafeEntryException, which no error handler can suppress.

Escaping symlinks are rejected by default with UnsafeEntryException. Use escapingSymlinkPolicy(ALLOW) for archives you produced yourself, or RELATIVIZE_ABSOLUTE to rewrite absolute targets under the output directory (targets that escape it are still rejected).

Decompression bombs are bounded by default: every archive extractor stops at 1,000,000 entries, and every extractor and Decompressor stops at an expansion ratio above 100 after the first MiB. Size limits are off. Use limits(…​), maxEntrySize, maxTotalSize and maxRatio to tune them, or limits(ExtractionLimits.noLimits()) for trusted input. Breaching a limit throws LimitExceededException, which an error handler cannot suppress. See Security.

try (TarArchiveExtractor extractor = TarArchiveExtractor.builder(Path.of("untrusted.tar.gz"))
        .maxEntries(10_000)
        .maxEntrySize(100L * 1024 * 1024)
        .maxTotalSize(1024L * 1024 * 1024)
        .build()) {
    extractor.extract(OUTPUT_DIR);
}

Key Goals

  • Simplicity: Offer an easy-to-understand API for common compression/decompression tasks.

  • Consistency: Provide a uniform interaction model across different compression algorithms.

  • Abstraction: Hide the complexities of individual compression library APIs from the end-user.

  • Extensibility: Facilitate the addition of support for new compression formats in the future.