Berkeley DB 4.8.30 ships a non-standard SHA-1 (and it changes with the machine's endianness)Hi Bitcointalk. I got curious about the encryption built into Berkeley DB, the DB_ENCRYPT feature you get with --enable-cryptography, and whether anyone in the early days ever used it to encrypt wallet files or other bitcoin data. BDB 4.8.30 is the version that stuck around the bitcoin world for years, so I went digging through its crypto code. I found something that would trip up anyone trying to decrypt that kind of data today, so I'm writing it up in case someone is sitting on an old encrypted database and can't get in.
TL;DRThe SHA-1 compiled into Berkeley DB 4.8.30 (hmac/sha1.c) is not SHA-1. A one-line C operator-precedence bug in the
blk0 macro corrupts the first 16 rounds, so
SHA1("abc") comes out as
c8afef54... instead of the textbook
a9993e36.... It round-trips fine inside BDB, so almost nobody noticed. Two things go wrong because of it:
- No standard tool (OpenSSL, Python hashlib, hashcat, any SHA-1 library) can reproduce BDB's HMAC/MAC or its AES key derivation. If you're trying to crack or recover a BDB-encrypted blob with a normal SHA-1, you will never match.
- The broken hash is architecture-dependent. A database encrypted with BDB's own DB_ENCRYPT on a little-endian box (x86) produces a different key and MAC than it would on a big-endian box, so the file can be permanently unreadable on the other architecture even with the correct password.
Worth saying up front: this is an interoperability and data-lock-in defect, not a cryptographic weakness, and it does not affect Bitcoin Core wallets (Core doesn't use BDB's built-in encryption, see the Scope section). It only bites code that used Berkeley DB's own
DB_ENCRYPT.
The bugIn
db-4.8.30/hmac/sha1.c around line 88:
Source :
https://download.oracle.com/berkeley-db/db-4.8.30.tar.gz#define blk0(i) is_bigendian ? block->l[i] : \
(block->l[i] = (rol(block->l[i],24)&0xFF00FF00) \
|(rol(block->l[i],8)&0x00FF00FF))
blk0() is meant to do one small thing. On a little-endian CPU it byte-swaps the i-th 32-bit message word into the big-endian order SHA-1 needs, and on a big-endian CPU it returns it unchanged. It gets used inside the first round macro R0:
#define R0(v,w,x,y,z,i) z+=((w&(x^y))^y)+blk0(i)+0x5A827999+rol(v,5); w=rol(w,30);
Because
blk0 has no parentheses around it, after macro expansion R0 becomes:
z += ((w&(x^y))^y) + is_bigendian ? block->l[i]
: (block->l[i]=SWAP) + 0x5A827999 + rol(v,5);
C precedence is
+ above
?: above
+=. With no parentheses the compiler reads it as:
z += ( ((w&(x^y))^y) + is_bigendian ) // the whole left side is the CONDITION
? ( block->l[i] ) // true branch
: ( (block->l[i]=SWAP) + 0x5A827999 + rol(v,5) ); // false branch
So in the 16 rounds that use blk0, on the branch that runs nearly every time, three things go missing:
- the nonlinear term ((w&(x^y))^y) gets swallowed into the condition and is never added to the state,
- the round constant 0x5A827999 is dropped,
- rol(v,5) is dropped,
- and the message word comes back without the byte-swap.
Only the first 16 rounds are affected, since R1 to R4 use
blk() which is parenthesized correctly, but trashing the first 16 rounds is enough to change the whole 80-round digest.
You can see it in one lineCompiling Berkeley DB's own unmodified
sha1.c on x86 (little-endian):
SHA1("abc") = c8afef5423a485a54f7ecba8bc778ce33e553f53 (standard: a9993e36...)
SHA1("") = a646ef710952b006be9498e2111045162233ef67 (standard: da39a3ee...)
HMAC(0x0b x20,"Hi There")= ac671d518a7334841d67a3c8e76600153401eb4e (RFC2202: b6173186...)
Put parentheses around the
blk0 ternary (the one-character fix) and the same file reproduces the exact FIPS-180-1 and RFC 2202 vectors. That proves the macro is the only thing wrong.
It is not weak, it is just a different hashBefore anyone panics, this does not make anything crackable. Over 300k random inputs the broken hash is statistically indistinguishable from real SHA-1:
out-bit bias (want 50%) avalanche /160 (want ~80)
BDB broken 50.00% (49.74..50.25) 79.99
standard SHA-1 50.00% (49.78..50.21) 79.99
Rounds 16 to 79 still mix everything properly, so what you get is a perfectly good 160-bit hash. It just isn't SHA-1, and nothing else computes it. No reduced security, full key entropy. The damage is interoperability, not security.
The part that actually loses dataReal SHA-1 gives the same answer on every machine. The broken one does not. Its dominant branch forwards the
raw machine-order word without normalizing it, so a little-endian build and a big-endian build end up computing two different functions:
SHA1("abc")
little-endian (x86) c8afef5423a485a54f7ecba8bc778ce33e553f53
big-endian (s390x) 3d11f750bf8d57b7b574d7cd95f8de44b52a9bbf
Over 20,000 random inputs the little-endian and big-endian digests disagreed on all 20,000. The control, real SHA-1, disagreed on 0 out of 2,000. To be sure the big-endian value wasn't just my model, I cross-compiled Berkeley DB's unmodified
sha1.c for IBM s390x and ran it under qemu. It produced exactly
3d11f750... for "abc", matching.
Why this makes data unreadable across machinesBDB's encryption builds everything out of this SHA-1:
- AES-128 key = BDB_SHA1( pw + "encryption and decryption key value magic" + pw )[:16]
- MAC key = BDB_SHA1( pw + "mac derivation key magic value" + pw )
- page MAC = HMAC-BDB_SHA1( mac_key, page )
Because the SHA-1 itself changes with the architecture, a database written with
DB_ENCRYPT on a little-endian server derives a different AES key and MAC than the same password would on a big-endian server. Open it on the other machine and the page MAC check fails, BDB reports it as a checksum or corruption error, and the plaintext decrypts to garbage, even with the correct password. It stays self-consistent only inside one endianness, which is exactly why it works fine until the day you move the file.
Scope, who is and isn't affected- Affected: anything that used Berkeley DB's own encryption (DB_ENCRYPT, builds with --enable-cryptography), whether directly through the C API or through some app or tool that turned it on. Especially if the data ever moved between a little-endian and a big-endian machine, or if you're trying to re-derive the key or MAC with a standard SHA-1.
- Not affected: Bitcoin Core wallet.dat. Core never used Berkeley DB's built-in encryption. Wallet secrets are encrypted at the application layer with its own AES-256-CBC and an OpenSSL based key derivation, so this code path is never touched. Ordinary unencrypted BDB files don't use SHA-1 for their page checksums either (they use a different function), so this only shows up when BDB's own crypto is switched on.
If you're stuck, how to get your data back- Open the database with a BDB build of the same endianness as the machine that wrote it. If it was written on x86, read it on an x86 box. Don't just add the parentheses and expect to read old data, because a corrected SHA-1 produces standard digests and will not match what's already stored.
- If you're writing a recovery or cracking tool, you have to reproduce BDB's broken SHA-1 exactly, matching the endianness of the machine that wrote the file, not OpenSSL or standard SHA-1. That one detail is the reason off-the-shelf tools never match a BDB-encrypted blob.
- If all you want is for new databases to be portable, parenthesize blk0 and rebuild, but migrate your old data out first with a matching-endianness build.
Reproduce it yourself- Grab db-4.8.30 and compile hmac/sha1.c on its own, then hash "abc". You get c8afef54... on x86. Add the parentheses to blk0 and you get a9993e36....
- Cross-compile the same file for a big-endian target like s390x under qemu and "abc" becomes 3d11f750....
I have a small test harness that does all of this side by side the unmodified file against the one-character fix, a round-by-round trace of where it goes wrong, the little vs big endian divergence test, and the s390x build. Happy to share it if anyone wants to check my work. If this bit you, reply with your architecture and how the data was originally made and I'll try to help. The fix is almost always to read it back on the same endianness it was written on.
Interesting no one found this at the time but again this would of been a non-standard thing to do but the options was there and I am sure some people may of used this option on mining servers back in the day.
Regards
SleepycatSwap
** Note I DID use AI to help format the post so it's clear where the issue is but the research and code was hand balled and tested.