Class FileHashService
- Namespace
- Archiver.Core.Services
- Assembly
- Archiver.Core.dll
T-F128: computes CRC-32/SHA-256 for the Explorer context menu's "Хеш-суми" submenu. A single
folder gets NanaZip-compatible combined DataSum (all file contents) and NamesSum (all file
names+paths+contents) values, via Archiver.Core.IO.HashDigestAccumulator — the algorithm of
7-Zip's HashCalc.cpp, checked live against the vendored 7za h by
FolderHashParityTests (T-F225). NamesSum includes one item per directory, the selected
folder itself included, hashed with an all-zero digest (7-Zip resets it before every item), so
the result does not depend on enumeration order. Remaining difference: symbolic links and
junctions are skipped here (T-F251), where 7-Zip follows them.
Files are hashed in parallel (ForEachAsync<TSource>(IEnumerable<TSource>, ParallelOptions, Func<TSource, CancellationToken, ValueTask>), up to ProcessorCount at once) — safe specifically because Add(ReadOnlySpan<byte>) is commutative, so combining DataSum/NamesSum in whatever order files finish hashing produces the exact same result as combining them sequentially (already relied on for the recursion-safety argument above).
T-F128 follow-up: a single large CRC-32 file is also hashed in parallel — the across-files parallelism above gives no benefit to a folder containing one huge file (or to a lone large file passed to ComputeAsync(IReadOnlyList<string>, HashAlgorithmKind, IProgress<ProgressReport>?, CancellationToken) directly). Files at or above Archiver.Core.Services.FileHashService.ParallelCrc32MinFileBytes are split into independently-hashed chunks and folded back together with Combine(uint, uint, long) — safe for the same reason cross-file combining is: CRC-32 combining is associative/order-preserving as long as chunks are folded back in their original byte order (unlike DataSum/NamesSum, chunk order here does matter, so chunks are combined sequentially by index after all finish, not as they complete).
public static class FileHashService
- Inheritance
-
FileHashService
- Inherited Members
Methods
ComputeAsync(IReadOnlyList<string>, HashAlgorithmKind, IProgress<ProgressReport>?, CancellationToken)
Hashes each of paths. A single-folder selection is routed to the
recursive DataSum/NamesSum path instead (see Folder); a mixed or
multi-item selection hashes each file independently in parallel.
public static Task<HashResult> ComputeAsync(IReadOnlyList<string> paths, HashAlgorithmKind algorithm, IProgress<ProgressReport>? progress, CancellationToken ct)
Parameters
pathsIReadOnlyList<string>algorithmHashAlgorithmKindprogressIProgress<ProgressReport>ctCancellationToken
Returns
ComputeStreamDigestAsync(Stream, HashAlgorithmKind, CancellationToken)
T-F09 follow-up (pakko h -si): hashes an arbitrary stream directly — a genuine single-pass
pipeline, unlike x/t/l's -si (which must stage stdin to a real
seekable temp file first, since ZIP central-directory reads and tar.exe's pre-scan can't
operate on a raw pipe). CRC-32/SHA-256 need no seeking, so no staging is needed here either —
bytes are read once, straight from source, with no intermediate file.
public static Task<string> ComputeStreamDigestAsync(Stream source, HashAlgorithmKind algorithm, CancellationToken ct)
Parameters
sourceStreamalgorithmHashAlgorithmKindctCancellationToken