Binary Diff & Patch
Compare binary artifacts and create deterministic forward or reverse copy-insert patches locked to verified SHA-256 identities.
Compare two builds and produce a patch you can verify
Text diff tools give up at the first non-printable byte. This one works on the bytes directly: load two artifacts, get a deterministic comparison, navigate the changes, and export a patch that is cryptographically bound to the file it was built from.
What the comparison gives you
The diff is expressed as a sequence of copy and insert operations, which is what makes it deterministic and what makes the patch format compact. Alongside that you get a similarity score, an estimate of how many bytes changed, an operation list you can step through, and a same-offset byte preview so you can see the before and after side by side at any position.
For most real comparisons the operation list is the useful view. A recompile of the same source produces a long copy, a short insert where a timestamp or build ID moved, and another long copy. A genuinely different binary produces a shredded operation list. You can usually tell those apart in seconds, before looking at a single byte.
Patches are locked to a hash
An exported patch uses the inventivehq-binary-patch/v1 JSON format and records the SHA-256 of both the source and the target. Applying it verifies the source hash first. Apply it to the file it was built from and you get a result whose hash matches the recorded target. Apply it to anything else and it is rejected rather than producing quiet corruption.
That property is the point of the tool. A patch that silently applies to the wrong input is worse than no patch, because the damage shows up later and somewhere else.
Patches run in both directions. A reverse patch takes the target back to the source, which is what you want when you are testing whether a change is the one causing a regression.
Things this is genuinely good for
- Confirming a vendor patch changed what the advisory says it changed. Diff the before and after binaries, then send the changed offsets to the Machine Code Disassembler and read the actual delta.
- Finding what differs between a working build and a broken one when the source tree claims they are identical.
- Checking whether two samples are the same family. A high similarity score with small scattered inserts usually means a reconfigured build rather than a rewrite.
- Distributing a small correction to a large file without shipping the whole file.
- Verifying that a file you received matches a file you have, where the answer needs to be more specific than "the hashes differ".
What byte-level diffing will not tell you
This release compares bytes. It does not understand executable sections, so inserting a few bytes near the start of a file can shift everything after it and produce a diff that looks enormous for a small logical change. It does not match functions across builds, does not diff semantically, and does not know that two different instruction encodings mean the same thing.
Section-aware and entropy-aware diffing, along with native semantic diffing and function matching, are deliberately scoped as later work rather than half-implemented here. For now, if a diff looks implausibly large, check whether an early insert shifted the alignment before concluding the files are unrelated.
Reading a diff that looks worse than it is
Two files compiled from identical source rarely produce an empty diff. Build timestamps, embedded paths, GUIDs, module IDs and signature blobs all move between builds, and each one shows up as a small insert. A handful of tiny operations scattered through a long copy is the normal signature of a rebuild, not evidence that something changed.
Alignment is the other common trap. Insert four bytes near the front of a file and every subsequent offset shifts, so a byte-level comparison reports a large fraction of the file as different even though almost nothing changed logically. When the similarity score looks implausibly low, step to the first operation: if it is an early insert followed by one enormous copy at a shifted offset, the files are closely related and the raw percentage is misleading.
Signed binaries deserve their own note. The signature blob sits at the end of the file and changes completely whenever anything before it changes, so it will always appear as a large trailing difference. Use the certificate table offset from the Executable & Object Inspector to identify that range and read the diff before it.
Limits and handoffs
Comparison is bounded in both input size and operation count, and imported patches are schema-validated and size-limited before anything is applied. Patched output lands in the Binary Lab workspace as a derived artifact that references its parent, so you can send the result straight into the Executable & Object Inspector or the disassembler without saving anything to disk.
You build the idea. I'll ship the product.
Productized MVP development for founders. 9 SaaS apps shipped — yours could be next, in 6 weeks. Secure by default.
Frequently Asked Questions
Common questions about the Binary Diff & Patch
A patch records the exact base and target size and SHA-256. Application refuses the wrong base, bounds-checks every operation, and hashes the reconstructed target.
Generate a reverse patch while both artifacts are loaded. It uses the former target as the hash-locked base and reconstructs the former base.
No. It uses the versioned inventivehq-binary-patch JSON format so operations and cryptographic identities remain inspectable and portable.
Explore More Tools
Continue with these related tools
⚠️ Security Notice
This tool is provided for educational and authorized security testing purposes only. Always ensure you have proper authorization before testing any systems or networks you do not own. Unauthorized access or security testing may be illegal in your jurisdiction. All processing happens client-side in your browser - no data is sent to our servers.