Bitcoin Forum
October 11, 2026, 11:11:13 PM *
News: Serious possible issue involving Ledger hardware wallets and CryptoBilis
 
   Home   Help Search Login Register More  
Pages: [1]
  Print  
Author Topic: DB 4.8.30 shipped a NON-STANDARD SHA-1  (Read 105 times)
SleepycatSwap (OP)
Newbie
*
Offline

Activity: 5
Merit: 14


View Profile
October 10, 2026, 09:51:00 AM
Merited by NotATether (5), Cricktor (2), vjudeu (1), stwenhao (1), athanred (1)
 #1

Berkeley DB 4.8.30 ships a non-standard SHA-1 (and it changes with the machine's endianness)

Hi Bitcointalk.  I got curious about the encryption built into Berkeley DB, the DB_ENCRYPT feature you get with --enable-cryptography, and whether anyone in the early days ever used it to encrypt wallet files or other bitcoin data. BDB 4.8.30 is the version that stuck around the bitcoin world for years, so I went digging through its crypto code. I found something that would trip up anyone trying to decrypt that kind of data today, so I'm writing it up in case someone is sitting on an old encrypted database and can't get in.

TL;DR
The SHA-1 compiled into Berkeley DB 4.8.30 (hmac/sha1.c) is not SHA-1. A one-line C operator-precedence bug in the blk0 macro corrupts the first 16 rounds, so SHA1("abc") comes out as c8afef54... instead of the textbook a9993e36.... It round-trips fine inside BDB, so almost nobody noticed. Two things go wrong because of it:
  • No standard tool (OpenSSL, Python hashlib, hashcat, any SHA-1 library) can reproduce BDB's HMAC/MAC or its AES key derivation. If you're trying to crack or recover a BDB-encrypted blob with a normal SHA-1, you will never match.
  • The broken hash is architecture-dependent. A database encrypted with BDB's own DB_ENCRYPT on a little-endian box (x86) produces a different key and MAC than it would on a big-endian box, so the file can be permanently unreadable on the other architecture even with the correct password.
Worth saying up front: this is an interoperability and data-lock-in defect, not a cryptographic weakness, and it does not affect Bitcoin Core wallets (Core doesn't use BDB's built-in encryption, see the Scope section). It only bites code that used Berkeley DB's own DB_ENCRYPT.

The bug
In db-4.8.30/hmac/sha1.c around line 88:

Source : https://download.oracle.com/berkeley-db/db-4.8.30.tar.gz

Code:
#define blk0(i) is_bigendian ? block->l[i] : \
    (block->l[i] = (rol(block->l[i],24)&0xFF00FF00) \
    |(rol(block->l[i],8)&0x00FF00FF))

blk0() is meant to do one small thing. On a little-endian CPU it byte-swaps the i-th 32-bit message word into the big-endian order SHA-1 needs, and on a big-endian CPU it returns it unchanged. It gets used inside the first round macro R0:

Code:
#define R0(v,w,x,y,z,i) z+=((w&(x^y))^y)+blk0(i)+0x5A827999+rol(v,5); w=rol(w,30);

Because blk0 has no parentheses around it, after macro expansion R0 becomes:

Code:
z += ((w&(x^y))^y) + is_bigendian ? block->l[i]
                                  : (block->l[i]=SWAP) + 0x5A827999 + rol(v,5);

C precedence is + above ?: above +=. With no parentheses the compiler reads it as:

Code:
z += ( ((w&(x^y))^y) + is_bigendian )      // the whole left side is the CONDITION
       ? ( block->l[i] )                    // true  branch
       : ( (block->l[i]=SWAP) + 0x5A827999 + rol(v,5) );   // false branch

So in the 16 rounds that use blk0, on the branch that runs nearly every time, three things go missing:
  • the nonlinear term ((w&(x^y))^y) gets swallowed into the condition and is never added to the state,
  • the round constant 0x5A827999 is dropped,
  • rol(v,5) is dropped,
  • and the message word comes back without the byte-swap.
Only the first 16 rounds are affected, since R1 to R4 use blk() which is parenthesized correctly, but trashing the first 16 rounds is enough to change the whole 80-round digest.

You can see it in one line
Compiling Berkeley DB's own unmodified sha1.c on x86 (little-endian):
Code:
SHA1("abc")              = c8afef5423a485a54f7ecba8bc778ce33e553f53   (standard: a9993e36...)
SHA1("")                 = a646ef710952b006be9498e2111045162233ef67   (standard: da39a3ee...)
HMAC(0x0b x20,"Hi There")= ac671d518a7334841d67a3c8e76600153401eb4e   (RFC2202: b6173186...)
Put parentheses around the blk0 ternary (the one-character fix) and the same file reproduces the exact FIPS-180-1 and RFC 2202 vectors. That proves the macro is the only thing wrong.

It is not weak, it is just a different hash
Before anyone panics, this does not make anything crackable. Over 300k random inputs the broken hash is statistically indistinguishable from real SHA-1:
Code:
                 out-bit bias (want 50%)      avalanche /160 (want ~80)
BDB broken       50.00% (49.74..50.25)        79.99
standard SHA-1   50.00% (49.78..50.21)        79.99
Rounds 16 to 79 still mix everything properly, so what you get is a perfectly good 160-bit hash. It just isn't SHA-1, and nothing else computes it. No reduced security, full key entropy. The damage is interoperability, not security.

The part that actually loses data
Real SHA-1 gives the same answer on every machine. The broken one does not. Its dominant branch forwards the raw machine-order word without normalizing it, so a little-endian build and a big-endian build end up computing two different functions:

Code:
                     SHA1("abc")
little-endian (x86)  c8afef5423a485a54f7ecba8bc778ce33e553f53
big-endian  (s390x)  3d11f750bf8d57b7b574d7cd95f8de44b52a9bbf

Over 20,000 random inputs the little-endian and big-endian digests disagreed on all 20,000. The control, real SHA-1, disagreed on 0 out of 2,000. To be sure the big-endian value wasn't just my model, I cross-compiled Berkeley DB's unmodified sha1.c for IBM s390x and ran it under qemu. It produced exactly 3d11f750... for "abc", matching.

Why this makes data unreadable across machines
BDB's encryption builds everything out of this SHA-1:
  • AES-128 key = BDB_SHA1( pw + "encryption and decryption key value magic" + pw )[:16]
  • MAC key = BDB_SHA1( pw + "mac derivation key magic value" + pw )
  • page MAC = HMAC-BDB_SHA1( mac_key, page )
Because the SHA-1 itself changes with the architecture, a database written with DB_ENCRYPT on a little-endian server derives a different AES key and MAC than the same password would on a big-endian server. Open it on the other machine and the page MAC check fails, BDB reports it as a checksum or corruption error, and the plaintext decrypts to garbage, even with the correct password. It stays self-consistent only inside one endianness, which is exactly why it works fine until the day you move the file.

Scope, who is and isn't affected
  • Affected: anything that used Berkeley DB's own encryption (DB_ENCRYPT, builds with --enable-cryptography), whether directly through the C API or through some app or tool that turned it on. Especially if the data ever moved between a little-endian and a big-endian machine, or if you're trying to re-derive the key or MAC with a standard SHA-1.
  • Not affected: Bitcoin Core wallet.dat. Core never used Berkeley DB's built-in encryption. Wallet secrets are encrypted at the application layer with its own AES-256-CBC and an OpenSSL based key derivation, so this code path is never touched. Ordinary unencrypted BDB files don't use SHA-1 for their page checksums either (they use a different function), so this only shows up when BDB's own crypto is switched on.

If you're stuck, how to get your data back
  • Open the database with a BDB build of the same endianness as the machine that wrote it. If it was written on x86, read it on an x86 box. Don't just add the parentheses and expect to read old data, because a corrected SHA-1 produces standard digests and will not match what's already stored.
  • If you're writing a recovery or cracking tool, you have to reproduce BDB's broken SHA-1 exactly, matching the endianness of the machine that wrote the file, not OpenSSL or standard SHA-1. That one detail is the reason off-the-shelf tools never match a BDB-encrypted blob.
  • If all you want is for new databases to be portable, parenthesize blk0 and rebuild, but migrate your old data out first with a matching-endianness build.

Reproduce it yourself
  • Grab db-4.8.30 and compile hmac/sha1.c on its own, then hash "abc". You get c8afef54... on x86. Add the parentheses to blk0 and you get a9993e36....
  • Cross-compile the same file for a big-endian target like s390x under qemu and "abc" becomes 3d11f750....

I have a small test harness that does all of this side by side the unmodified file against the one-character fix, a round-by-round trace of where it goes wrong, the little vs big endian divergence test, and the s390x build. Happy to share it if anyone wants to check my work. If this bit you, reply with your architecture and how the data was originally made and I'll try to help. The fix is almost always to read it back on the same endianness it was written on.

Interesting no one found this at the time but again this would of been a non-standard thing to do but the options was there and I am sure some people may of used this option on mining servers back in the day.

Regards

SleepycatSwap

** Note I DID use AI to help format the post so it's clear where the issue is but the research and code was hand balled and tested.
NotATether
Legendary
*
Offline

Activity: 2478
Merit: 10390


┻┻ ︵㇏(°□°㇏)


View Profile WWW
October 10, 2026, 10:44:57 AM
 #2

I don't think people should be using such cryptographic primitives provided by database in the first place.

BDB has also never been audited like eg, OpenSSL.

So it would be very foolish to make applications depend on this specific feature.

▄▄████████████████████▄▄
▄███████▀▀██████▀▀███████▄
█████▀██████████████▀█████
████████▄▄██████▄▄████▀███

██████████████████████████
██▄▄██████████████▄▄██████
██▀▀██████████████████▄▄██
██████▀▀██████████████▀▀██
██████████████████████████
███▄████▀▀██████▀▀████████
█████▄██████████████▄█████
▀███████▄▄██████▄▄███████▀
▀▀████████████████████▀▀
 
 DΞX.fo 
▄▄██████
█████████
██████████
██████████
██████████
█████████
▀▀██████

▄███████
▄██████████
████████████
█████████████
█████████████
|
▄▄█
▄████▀
▄███▀█▄
▄██▀█▄██
█████▀▀█
████████
████████
▀██▄████
▄████▄▄█
▄█████▀███
▄█████▀████▀
█████▀███████
▀██▀█████████
|..BTC......XMR...
..USDT.....LTC...
....Fees  0.8%.....
athanred
Full Member
***
Offline

Activity: 191
Merit: 345


View Profile
Today at 02:36:13 AM
Merited by NotATether (3)
 #3

Quote
in case someone is sitting on an old encrypted database and can't get in
Seems to be unlikely, but it is a nice finding, if it is true. I will try to verify it.

Hashes reproduced by me so far:
Code:
SHA1("abc")=a9993e364706816aba3e25717850c26c9cd0d89d
SHA1("")=da39a3ee5e6b4b0d3255bfef95601890afd80709
HMAC-SHA1(0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b0b,"Hi There")=b617318655057264e28bc0b6fb378c8ef146be00
I will try to verify bugged ones later, but they seems to be correct.

I think no users are affected, however, maybe it will help uncover the mystery behind the low x-value in secp256k1 generator, or some elliptic curve seeds. If they used some old code, then SHA1("Base point"||some_data) can be non-standard, which could explain, why nobody reproduced these seeds so far.
SleepycatSwap (OP)
Newbie
*
Offline

Activity: 5
Merit: 14


View Profile
Today at 11:55:04 AM
Merited by NotATether (3), vjudeu (1)
 #4

Thankis but did you run those with a normal SHA1 though, hashlib or openssl, so of course they match the textbook. That was never the part I said was broken.

The broken one is the SHA1 that Berkeley DB compiles out of its own hmac/sha1.c.

That code never talks to anything outside BDB, so it never had to agree with a real SHA1, and that's exactly why it went unnoticed this long. Run "abc" through that file instead of openssl and you get a different digest:

Code:
                           BDB 4.8.30 (own sha1.c)                   standard SHA1
SHA1("abc")                c8afef5423a485a54f7ecba8bc778ce33e553f53  a9993e364706816aba3e25717850c26c9cd0d89d
SHA1("")                   a646ef710952b006be9498e2111045162233ef67  da39a3ee5e6b4b0d3255bfef95601890afd80709
HMAC(0x0b x20,"Hi There")  ac671d518a7334841d67a3c8e76600153401eb4e  b617318655057264e28bc0b6fb378c8ef146be00

Easiest way to see it: pull hmac/sha1.c out of db-4.8.30, compile it on its own, hash "abc", and you get c8afef54. Then put parentheses around the blk0 ternary and rebuild.

The same file now prints a9993e36, da39a3ee and the right RFC2202 HMAC. One change, nothing else touched, so it's the macro.

The why, if you want it: blk0 isn't parenthesized, and + binds tighter than ?: which binds tighter than +=, so R0 actually compiles as:

Code:
z += ( ((w&(x^y))^y) + is_bigendian ) ? block->l[i] : (block->l[i]=SWAP) + 0x5A827999 + rol(v,5);

So on the branch that runs basically every round, the F term, the 0x5A827999 and the rol(v,5) all just vanish and the word never gets byte-swapped. Only the 16 R0 rounds hit it, R1 through R4 use blk() which is fine, but that's plenty to wreck the digest.

@NotATether, no argument, nobody should be leaning on BDB's crypto. But some people did. DB_ENCRYPT was the only at-rest option on server builds before Core added wallet encryption in 0.4, and those files are still out there. That's who this is for, and it's why a standard SHA1 never matches them.

Got the harness if anyone wants to run it themselves, the unmodified file vs the one-char fix, a round trace, LE vs BE, and the s390x build.
SleepycatSwap (OP)
Newbie
*
Offline

Activity: 5
Merit: 14


View Profile
Today at 05:25:46 PM
 #5

Quote
in case someone is sitting on an old encrypted database and can't get in
I think no users are affected, however, maybe it will help uncover the mystery behind the low x-value in secp256k1 generator, or some elliptic curve seeds. If they used some old code, then SHA1("Base point"||some_data) can be non-standard, which could explain, why nobody reproduced these seeds so far.

I don't think it reaches the secp256k1 seeds. This bug is entirely inside Berkeley DB's own sha1.c and nothing in the curve generation ever calls that file, so the low x-value and the unreproduced seeds are a separate rabbit hole. Different code path, different story.

Where I do think people got caught is the early users. If someone turned on the database encryption back then and ran the tools against it, everything looked fine for as long as they stayed on the same build. The broken digest was at least consistent with itself, so it opened and closed without complaint. The trouble only shows up later, when they move to a new machine or a rebuilt binary and the hash no longer comes out the same. That is when the file suddenly won't open, and from the outside it looks like a corrupt wallet rather than a hashing bug.

It's an unseen failure for exactly that reason. Nobody would have noticed it in normal day to day use, so there could easily be old miners or early users sitting on a file they can't get back into and no idea why.

Worth remembering the options at the time too. If you wanted to protect a wallet the usual route was OpenSSL or PGP on the file yourself, but the DB that shipped with Bitcoin had encryption built in, so we can't rule out that some people leaned on that instead and are stuck now.

Scope wise it's narrow. It only touches builds with crypto enabled where the SHA-1 or HMAC functions actually got used. If that never happened on your setup there's nothing here for you. But for the handful it did happen to, this is probably the reason the digest breaks the moment they change systems.
TheButterZone
Legendary
*
Offline

Activity: 3290
Merit: 1146


RIP Mommy


View Profile WWW
Today at 05:37:31 PM
 #6

So now it's a matter of trying to programmatically find everyone who reported wallets from this time being broken, and then actually managing to make contact & direct them to this thread?

Some may never come back to wherever they posted, or even have changed/lost their email address, so even notifications they enabled would fail.

SleepycatSwap (OP)
Newbie
*
Offline

Activity: 5
Merit: 14


View Profile
Today at 05:54:11 PM
 #7

So now it's a matter of trying to programmatically find everyone who reported wallets from this time being broken, and then actually managing to make contact & direct them to this thread?

Some may never come back to wherever they posted, or even have changed/lost their email address, so even notifications they enabled would fail.

In a perfect world, I posted it just on the off chance there is still someone sitting on such a DB or wallet from the period that might think it's broken but it's not.

Historical bug, but matters if bitcoin was built with those flags enabled on the DB so best to post it on the off-chance it can help someone.
Pages: [1]
  Print  
 
Jump to:  

Powered by MySQL Powered by PHP Powered by SMF 1.1.19 | SMF © 2006-2009, Simple Machines Valid XHTML 1.0! Valid CSS!