Reading
ArchiveExtractor.stream() returns the archive’s entries in order as ArchiveItem`s, after `stripComponents, filter and the entry limit apply. item.content() reads the entry’s bytes, metered against maxEntrySize, maxTotalSize and maxRatio.
try (var extractor = TarArchiveExtractor.builder(archive).build()) {
return extractor.stream()
.filter(item -> item.entry().name().equals("config/app.properties"))
.findFirst()
.map(item -> {
try {
return new String(item.content().readAllBytes(), StandardCharsets.UTF_8);
} catch (IOException e) {
throw new UncheckedIOException(e);
}
});
} catch (UncheckedIOException e) {
throw e.getCause();
}
An item’s content is readable only while it is current. Collecting items first and reading later fails:
try (var extractor = TarArchiveExtractor.builder(archive).build()) {
var items = extractor.stream().toList(); // advances past every entry
items.getFirst().content(); // throws IllegalStateException: entry ... is no longer current
}
The stream is sequential and single-use: a parallel stream fails with IllegalStateException when consumed, and a second stream() or extract() call throws it too. stream() surfaces parser failures, including corrupt input, as UncheckedIOException wrapping the IOException.
Decompressing a stream
Decompressor.inputStream() returns the decompressed bytes of one compressed stream, metered against the same ratio and size limits. Closing it closes the Decompressor. Combine it with any archive or parser that reads an InputStream.
Writing
ArchiveCreator.addAll writes a stream of EntrySource`s: files, directories and symlinks from disk (`EntrySource.of), memory (EntrySource.file) or another archive (ArchiveItem.toSource).
try (var creator = TarArchiveCreator.builder(archive)
.compression(Compression.gzip())
.build();
Stream<Path> paths = Files.walk(dir)) {
creator.addAll(
paths.skip(1).filter(p -> !p.toString().endsWith(".tmp")).map(p -> {
try {
return EntrySource.of(dir, p);
} catch (IOException e) {
throw new UncheckedIOException(e);
}
}));
}
tar, ar and cpio write each file’s size before its content. A source without a size fails with IllegalArgumentException; spool it explicitly:
var unsized = new EntrySource.File(
"download.bin", 0, FileTime.from(Instant.now()), OptionalLong.empty(), () -> download);
try (var creator = TarArchiveCreator.builder(archive)
.compression(Compression.gzip())
.build()) {
creator.add(EntrySource.buffered(unsized, tempDir));
}
EntrySource.buffered closes the source stream it reads.
Repacking
try (var extractor = TarArchiveExtractor.builder(tarGz).build();
var creator = ZipArchiveCreator.builder(zip).build()) {
creator.addAll(extractor.stream().map(ArchiveItem::toSource));
}
The creator checks entry names on write, so repacking an archive with ../ names fails with UnsafeEntryException.
Format capabilities
Verified by the contract suite on every build:
| Format | directories | modes | symlinks | last modified | requires size | stream input | stream output | random access input |
|---|---|---|---|---|---|---|---|---|
tar |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.gz |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.bz2 |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.xz |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.lzma |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.lz4 |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
tar.zst |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
ar |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
||
cpio |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
|
zip |
✓ |
✓ |
✓ |
✓ |
✓ |
✓ |
||
zip-streaming |
✓ |
✓ |
✓ |
|||||
7z |
✓ |
✓ |
✓ |
✓ |
✓ |
-
directories, modes, symlinks, last modified: the format preserves them on a round trip.
-
requires size: sizes must be known before content, so a source without a size needs
EntrySource.buffered. -
stream input, stream output: the format reads from an
InputStreamor writes to anOutputStream. -
random access input: reads from the channel’s start regardless of its position; other formats read from the current position.
-
Zip’s stream input is the zip-streaming row.