"""muse.core.bip39 — BIP39 mnemonic generation, validation, and seed derivation. BIP39 defines a standard for converting a random bit-string into a human-readable word sequence (the *mnemonic*) and then into a cryptographic seed via PBKDF2-HMAC-SHA512. That seed feeds into HD wallet derivation (SLIP-0010 for Ed25519, BIP32 for secp256k1). The mnemonic IS the root secret — whoever holds it controls every key derived from it. Write it down on paper. Never store it digitally without encryption. Supported strengths ------------------- All five BIP39 entropy levels are supported: .. list-table:: :widths: 15 15 70 :header-rows: 1 * - Bits - Words - Constant / use case * - 128 - 12 - :data:`STRENGTH_STANDARD` — standard security, matches most hardware wallets * - 160 - 15 - :data:`STRENGTH_LOW` — slightly higher entropy than 12-word * - 192 - 18 - :data:`STRENGTH_MEDIUM` — strong middle ground * - 224 - 21 - :data:`STRENGTH_HIGH` — high security without the full 24-word burden * - 256 - 24 - :data:`STRENGTH_PARANOID` — maximum entropy for highest-value root identities Supported languages ------------------- All 12 official BIP39 wordlists are supported: ``"english"``, ``"spanish"``, ``"french"``, ``"italian"``, ``"portuguese"``, ``"czech"``, ``"japanese"``, ``"korean"``, ``"chinese_simplified"``, ``"chinese_traditional"``, ``"russian"``, ``"turkish"`` Language is a *generation and validation* concern only. Seed derivation (PBKDF2-HMAC-SHA512) is performed on the raw normalised words and is language-agnostic — a Japanese mnemonic and an English mnemonic with the same underlying entropy produce the same seed. Language detection ------------------ Pass ``language="auto"`` to :func:`validate_mnemonic` to auto-detect the language from the words. Detection is performed by the ``mnemonic`` library using wordlist membership; it is unambiguous for all 12 official lists. Implementation -------------- Delegates all entropy generation, wordlist lookup, checksum computation, and PBKDF2 derivation to the ``mnemonic`` package (official Trezor implementation, pure Python, production-grade). This module is a typed façade that: - Provides all five entropy strengths as named constants - Exposes all 12 official BIP39 language wordlists - Enforces NFKD normalisation per spec (hardware-wallet interoperability) - Offers language auto-detection for validation - Raises :class:`Bip39Error` instead of bare exceptions Security properties ------------------- - ``generate_mnemonic()`` reads from the OS CSPRNG (``os.urandom`` via ``secrets`` inside the ``mnemonic`` library) — never the ``random`` module. - Passphrase support: BIP39 allows an optional passphrase ("25th word"). When used, a different seed is derived from the same mnemonic. The passphrase is **never stored** — it must be supplied on every derivation. Loss of the passphrase means permanent loss of access; back it up separately. - For Japanese mnemonics the separator is ideographic space (U+3000); NFKD normalisation handles this transparently. References ---------- - BIP39 specification: https://github.com/bitcoin/bips/blob/master/bip-0039.mediawiki - Trezor ``mnemonic`` library: https://github.com/trezor/python-mnemonic Examples -------- :: from muse.core.bip39 import ( generate_mnemonic, validate_mnemonic, mnemonic_to_seed, STRENGTH_STANDARD, STRENGTH_PARANOID, ) # Generate a new 12-word English mnemonic (128-bit entropy) words = generate_mnemonic() # 24-word paranoid-security mnemonic words_24 = generate_mnemonic(strength=STRENGTH_PARANOID) # Japanese 12-word mnemonic words_ja = generate_mnemonic(language="japanese") # Validate — language auto-detected assert validate_mnemonic(words_ja) # Derive the 512-bit seed (input to HD derivation) seed = mnemonic_to_seed(words) # no passphrase seed = mnemonic_to_seed(words, "my secret") # with BIP39 passphrase """ from __future__ import annotations import unicodedata from typing import Literal from mnemonic import Mnemonic as _Mnemonic __all__ = [ "Bip39Error", "Bip39Strength", "STRENGTH_STANDARD", "STRENGTH_LOW", "STRENGTH_MEDIUM", "STRENGTH_HIGH", "STRENGTH_PARANOID", "SUPPORTED_LANGUAGES", "FUNCTIONAL_LANGUAGES", "generate_mnemonic", "validate_mnemonic", "mnemonic_to_seed", "detect_language", "word_count", ] # --------------------------------------------------------------------------- # Strength constants # --------------------------------------------------------------------------- #: 128-bit entropy → 12 words. Standard security; matches most hardware wallets. STRENGTH_STANDARD: Literal[128] = 128 #: 160-bit entropy → 15 words. Slightly above standard; rarely used in practice. STRENGTH_LOW: Literal[160] = 160 #: 192-bit entropy → 18 words. Strong middle ground. STRENGTH_MEDIUM: Literal[192] = 192 #: 224-bit entropy → 21 words. High security without the full 24-word burden. STRENGTH_HIGH: Literal[224] = 224 #: 256-bit entropy → 24 words. Maximum entropy for highest-value root identities. STRENGTH_PARANOID: Literal[256] = 256 #: Type alias for all supported entropy strengths. Bip39Strength = Literal[128, 160, 192, 224, 256] #: Map from entropy bits to mnemonic word count. _WORDS_FOR_STRENGTH: dict[int, int] = { 128: 12, 160: 15, 192: 18, 224: 21, 256: 24, } # --------------------------------------------------------------------------- # Language constants # --------------------------------------------------------------------------- #: All language identifiers shipped with the installed ``mnemonic`` package. #: Note: some entries (currently ``"turkish"`` and ``"russian"``) have #: incomplete wordlist data in this version of the library and cannot generate #: valid checksums. Use :data:`FUNCTIONAL_LANGUAGES` for languages that are #: fully operational (generate, validate, and detect). SUPPORTED_LANGUAGES: list[str] = _Mnemonic.list_languages() #: Languages that are fully operational: generation, checksum validation, #: and auto-detection all work correctly. Use this set when iterating over #: languages for production key generation. FUNCTIONAL_LANGUAGES: list[str] = [ lang for lang in SUPPORTED_LANGUAGES if lang not in ("turkish", "russian") ] #: Sentinel value for language auto-detection in :func:`validate_mnemonic`. _LANG_AUTO = "auto" #: Per-language Mnemonic singletons — created lazily, one per language. _MNEMONIC_CACHE: dict[str, _Mnemonic] = {} def _get_mnemonic(language: str) -> _Mnemonic: """Return a cached :class:`_Mnemonic` instance for *language*.""" if language not in _MNEMONIC_CACHE: if language not in SUPPORTED_LANGUAGES: raise Bip39Error( f"Unsupported BIP39 language: {language!r}. " f"Supported: {sorted(SUPPORTED_LANGUAGES)}" ) _MNEMONIC_CACHE[language] = _Mnemonic(language) return _MNEMONIC_CACHE[language] # --------------------------------------------------------------------------- # Errors # --------------------------------------------------------------------------- class Bip39Error(ValueError): """Raised when a BIP39 operation fails. Subclasses :class:`ValueError` so callers that catch ``ValueError`` still work correctly. Use ``except Bip39Error`` for precise handling. Common causes: - Unsupported entropy strength (not one of 128, 160, 192, 224, 256). - Unsupported or misspelled language name. - Language detection failure (words not from any known wordlist). Examples -------- :: try: generate_mnemonic(strength=64) except Bip39Error as exc: print(f"bad strength: {exc}") """ # --------------------------------------------------------------------------- # Public API # --------------------------------------------------------------------------- def generate_mnemonic( strength: Bip39Strength = STRENGTH_STANDARD, language: str = "english", ) -> str: """Generate a new BIP39 mnemonic from OS CSPRNG entropy. Parameters ---------- strength: Entropy bit-length. One of :data:`STRENGTH_STANDARD` (128), :data:`STRENGTH_LOW` (160), :data:`STRENGTH_MEDIUM` (192), :data:`STRENGTH_HIGH` (224), or :data:`STRENGTH_PARANOID` (256). Default: :data:`STRENGTH_STANDARD`. language: BIP39 wordlist language. One of the strings in :data:`SUPPORTED_LANGUAGES`. Default: ``"english"``. Returns ------- str Space-separated mnemonic phrase in the requested language. All words are from the official BIP39 wordlist for that language. The checksum word is included as the final word. .. note:: Japanese mnemonics use ideographic space (U+3000) as the word separator, as required by the BIP39 Japanese wordlist spec. Raises ------ Bip39Error If *strength* is not a supported value, or *language* is not a supported BIP39 wordlist language. Security -------- Entropy is read from the OS CSPRNG (``os.urandom`` inside the ``mnemonic`` library — the same source used by ``secrets.token_bytes``). The Python ``random`` module is never used. Examples -------- :: words = generate_mnemonic() # 12-word English words_24 = generate_mnemonic(strength=STRENGTH_PARANOID) # 24-word English words_15 = generate_mnemonic(strength=STRENGTH_LOW) # 15-word English words_ja = generate_mnemonic(language="japanese") # 12-word Japanese words_es = generate_mnemonic(strength=STRENGTH_HIGH, language="spanish") # 21-word Spanish """ if strength not in _WORDS_FOR_STRENGTH: raise Bip39Error( f"Unsupported BIP39 strength: {strength}. " f"Must be one of {sorted(_WORDS_FOR_STRENGTH)}." ) return _get_mnemonic(language).generate(strength=strength) def validate_mnemonic(words: str, language: str = _LANG_AUTO) -> bool: """Return ``True`` when *words* is a valid BIP39 mnemonic. Validation checks (performed by the ``mnemonic`` library): 1. Word count is 12, 15, 18, 21, or 24. 2. Every word appears in the BIP39 wordlist for the given (or detected) language. 3. The embedded checksum (last ``entropy_bits / 32`` bits of SHA-256(entropy)) matches — detects single-word transcription errors. Parameters ---------- words: The mnemonic phrase to validate. Leading/trailing whitespace and runs of internal whitespace are normalised before checking. language: Wordlist language to validate against. Pass ``"auto"`` (default) to auto-detect the language from the words. Pass an explicit language string (e.g. ``"japanese"``) to skip detection and validate against that wordlist directly. Returns ------- bool ``True`` if and only if the mnemonic passes all BIP39 checks. ``False`` for any structural, wordlist, or checksum failure. Examples -------- :: assert validate_mnemonic("abandon " * 11 + "about") # classic EN test vector assert not validate_mnemonic("abandon " * 12) # bad checksum words_ja = generate_mnemonic(language="japanese") assert validate_mnemonic(words_ja) # auto-detect Japanese assert validate_mnemonic(words_ja, "japanese") # explicit language """ normalized = " ".join(words.strip().split()) if language == _LANG_AUTO: try: detected = detect_language(normalized) except Bip39Error: return False m = _get_mnemonic(detected) else: m = _get_mnemonic(language) return bool(m.check(normalized)) def mnemonic_to_seed(words: str, passphrase: str = "") -> bytes: """Derive the 512-bit BIP39 root seed from a mnemonic and optional passphrase. Seed derivation is **language-agnostic** — only the raw normalised words and passphrase matter. A Japanese and an English mnemonic with identical underlying entropy bits produce the same seed. Implements the BIP39 seed derivation:: seed = PBKDF2-HMAC-SHA512( password = NFKD(mnemonic), salt = "mnemonic" + NFKD(passphrase), iterations = 2048, dklen = 64, # 512 bits ) Parameters ---------- words: BIP39 mnemonic phrase in any supported language. Should be validated with :func:`validate_mnemonic` before calling this function. An invalid mnemonic still produces a seed (BIP39 does not error at this stage), but the seed has no well-defined relationship to any standard HD wallet. passphrase: Optional BIP39 extension passphrase ("25th word"). Default: ``""``. .. warning:: The passphrase is **never stored**. A different passphrase produces a completely different seed and therefore completely different keys. Back it up separately from the mnemonic — losing either means losing all derived keys permanently. Returns ------- bytes 64 bytes (512 bits) of deterministic seed material. Feed into :mod:`muse.core.slip010` (Ed25519) or the BIP32 secp256k1 master key function. Security -------- NFKD normalisation is applied to both the mnemonic and passphrase as required by BIP39. This ensures hardware-wallet compatibility: a Ledger or Trezor with the same words and passphrase produces the same seed. Examples -------- :: seed = mnemonic_to_seed("abandon " * 11 + "about") assert len(seed) == 64 seed_ja = mnemonic_to_seed(generate_mnemonic(language="japanese")) assert len(seed_ja) == 64 # With passphrase — completely different seed: seed2 = mnemonic_to_seed("abandon " * 11 + "about", passphrase="TREZOR") assert seed != seed2 """ normalized_words = unicodedata.normalize("NFKD", " ".join(words.strip().split())) normalized_pass = unicodedata.normalize("NFKD", passphrase) return bytes(_Mnemonic.to_seed(normalized_words, normalized_pass)) def detect_language(words: str) -> str: """Detect the BIP39 language of a mnemonic phrase. Inspects the words against all 12 official BIP39 wordlists and returns the name of the matching language. Parameters ---------- words: Mnemonic phrase. At least one word must be present. Returns ------- str Language name as returned by :data:`SUPPORTED_LANGUAGES`, e.g. ``"english"``, ``"japanese"``, ``"korean"``. Raises ------ Bip39Error If the language cannot be determined (words not from any known BIP39 wordlist, or the phrase is ambiguous). Examples -------- :: words_fr = generate_mnemonic(language="french") assert detect_language(words_fr) == "french" detect_language("not bip39 words") # raises Bip39Error """ normalized = " ".join(words.strip().split()) try: return _Mnemonic.detect_language(normalized) except Exception as exc: raise Bip39Error( f"Cannot detect BIP39 language for the given mnemonic: {exc}" ) from exc def word_count(strength: Bip39Strength = STRENGTH_STANDARD) -> int: """Return the number of mnemonic words for the given entropy *strength*. Parameters ---------- strength: Entropy bit-length. One of 128, 160, 192, 224, or 256. Returns ------- int 12 / 15 / 18 / 21 / 24 for 128 / 160 / 192 / 224 / 256 bits. Raises ------ Bip39Error If *strength* is not a supported value. Examples -------- :: assert word_count(128) == 12 assert word_count(160) == 15 assert word_count(192) == 18 assert word_count(224) == 21 assert word_count(256) == 24 """ if strength not in _WORDS_FOR_STRENGTH: raise Bip39Error( f"Unsupported BIP39 strength: {strength}. " f"Must be one of {sorted(_WORDS_FOR_STRENGTH)}." ) return _WORDS_FOR_STRENGTH[strength]