Skip to content

validation: cap mainchain RPC timeout during peg-in validation (+ -mainchainrpctimeout=0 help text fix) - #1608

Open
GracedEternalKingCabbageMan wants to merge 2 commits into
ElementsProject:masterfrom
GracedEternalKingCabbageMan:validatepegin-rpc-timeout-hardening
Open

GracedEternalKingCabbageMan wants to merge 2 commits into
ElementsProject:masterfrom
GracedEternalKingCabbageMan:validatepegin-rpc-timeout-hardening

Conversation

@GracedEternalKingCabbageMan

Copy link
Copy Markdown

validation: cap mainchain RPC timeout during peg-in validation

This PR contains two commits:

  1. rpc: correct the -mainchainrpctimeout=0 help text — a one-line documentation fix.
  2. validation: cap mainchain RPC timeout during peg-in validation — the behavior change.

Problem

Peg-in validation (-validatepegin, default-on for chains with a federated peg such as Liquid)
issues synchronous JSON-RPC calls to the configured mainchain daemon, from code paths that
hold cs_main (and, for mempool acceptance, mempool.cs as well):

Path Call Lock context
Block connection: Chainstate::ConnectTip → CheckPeginRipeness getblockheader (per peg-in input) AssertLockHeld(cs_main)
Consensus::CheckTxInputs, called from both block connection (ConnectBlock) and mempool acceptance (MemPoolAccept::PreChecks) getblockheader (per peg-in input) cs_main (+ mempool.cs at acceptance)
Mempool acceptance: MemPoolAccept::PreChecks → CheckPeginSubsidyAndMinimum getrawtransaction (per peg-in input) EXCLUSIVE_LOCKS_REQUIRED(cs_main, m_pool.cs)
Mempool epoch ejection: CTxMemPool::removeForBlock getblockheader (per peg-in input) connect path

All of these go through CallMainChainRPC, which applies -mainchainrpctimeout (default
900 seconds) to each call, with a fresh blocking TCP connection per call and no caching of
results.

If the local mainchain daemon accepts TCP connections but stops answering requests (it is itself
stalled, overloaded, or otherwise wedged), every one of these calls blocks for up to 15 minutes
per peg-in input while the locks are held — freezing block validation, mempool acceptance, and
most RPC for the duration. ConnectTip even logs that the chain "will not grow until this is
remedied", so this coupling is by design; what is missing is a bound on how long a single
unresponsive-daemon episode can stall validation per attempt.

Notably, failing fast is already the safe semantic on every one of these paths:

  • IsConfirmedBitcoinBlock catches the connection failure and returns false — "not yet
    confirmed, retry later" (at block connection this is the deliberate stall-and-retry machinery;
    in the mempool the transaction simply stays).
  • The getrawtransaction call in CheckPeginSubsidyAndMinimum propagates the same
    CConnectionFailed it would hit after 15 minutes today, into the same retryable
    mempool-rejection path ("pegin-subsidy-mainchain-error").

So blocking for the full 900 s buys nothing on these paths: after 30 s the code lands in exactly
the same state it reaches after 900 s today.

Change

  • CallMainChainRPC gains an optional timeout parameter. Negative (the default) preserves the
    current behavior for every existing caller: read -mainchainrpctimeout.
  • New GetValidationRPCTimeout() returns -mainchainrpctimeout clamped into
    [1, MAX_VALIDATION_RPC_TIMEOUT] (new constant, 30 seconds).
  • The clamped timeout is used at exactly the two validation-path call sites:
    IsConfirmedBitcoinBlock (getblockheader) and the getrawtransaction call in
    CheckPeginSubsidyAndMinimum. IsConfirmedBitcoinBlock also serves the advisory mature
    field of createrawpegin, which fails faster against a hung daemon — harmless and desirable.
  • The startup sanity check (MainchainRPCCheck in init.cpp) and all other CallMainChainRPC
    users (wallet RPCs) are unchanged and keep the full -mainchainrpctimeout.

30 s is generous for the intended deployment (a co-located, trusted bitcoind answering
getblockheader/getrawtransaction) and bounds a stall to well under one Liquid block interval
per input instead of 15 minutes.

Help text fix (first commit)

The -mainchainrpctimeout help text claims "0 for no timeout". That is not what happens: with
the pinned libevent 2.1.12, a timeout of 0 is treated as unset and libevent substitutes its own
internal defaults (45 s connect, 50 s read/write — HTTP_CONNECT_TIMEOUT/HTTP_READ_TIMEOUT/
HTTP_WRITE_TIMEOUT in http-internal.h, applied in http.c). There is no way to disable the
timeout through this option. The help text now states the actual behavior (and, in the second
commit, documents the validation cap).

Backward compatibility

  • Operators who set -mainchainrpctimeout to a value ≤ 30 see no change in validation behavior;
    values above 30 (including the 900 s default) are capped at 30 s for the two validation-path
    call sites only.
  • All other mainchain RPC consumers keep the configured timeout.
  • No new options, no changes to consensus rules, mempool policy, or RPC interfaces.
  • Failure semantics are unchanged: the same false-retry and exception paths are taken, only
    sooner.

Test

Adds mainchainrpc_tests.cpp covering GetValidationRPCTimeout: default clamped to the cap,
values below the cap honored, values above the cap clamped, boundary value honored, and 0
clamped to 1 (since 0 never meant "no timeout" anyway). A functional harness for a
hung-but-connected daemon was considered out of scope for this PR; happy to add one if reviewers
want end-to-end coverage of the stall path.

Notes for reviewers

  • Cap value: 30 s was chosen as generous for the intended deployment (a co-located, trusted
    bitcoind) while bounding a stall to well under one Liquid block interval; 60 s would work
    just as well if preferred.
  • Clamp vs. knob: a separate configurable validation-path timeout was considered and
    rejected to keep the option surface minimal; the patch is structured so that swapping the
    constant for an argument is a two-line change if maintainers prefer that.
  • Backport: the touched logic is identical on the 23.3.x maintenance line; the patch does
    not apply there verbatim (autotools vs. CMake test registration, minor context drift) but is a
    trivial manual re-application if wanted.

The help text claims a value of 0 disables the timeout. It does not:
with the pinned libevent 2.1.12, a timeout of 0 is treated as unset and
libevent substitutes its own internal defaults (45s connect, 50s
read/write). There is no way to disable the timeout via this option.
Peg-in validation (-validatepegin, default on for chains with a federated
peg) issues synchronous mainchain RPC calls while holding cs_main: during
block connection (ConnectTip via CheckPeginRipeness; ConnectBlock via
Consensus::CheckTxInputs), during mempool acceptance (PreChecks via
CheckPeginSubsidyAndMinimum), and during mempool epoch ejection
(removeForBlock). Each call uses the full -mainchainrpctimeout (default
900s) with a fresh blocking TCP connection and no caching, so a mainchain
daemon that accepts connections but stops responding can stall block
validation and mempool acceptance for up to 15 minutes per peg-in input.

IsConfirmedBitcoinBlock treats RPC failure as 'not confirmed, retry
later', and the getrawtransaction call in CheckPeginSubsidyAndMinimum
fails the transaction into the same retryable mempool-rejection path, so
blocking for the full timeout has no benefit on these paths.

Give CallMainChainRPC an optional timeout parameter (negative preserves
the existing behavior for all other callers), add
GetValidationRPCTimeout() clamping -mainchainrpctimeout into
[1, MAX_VALIDATION_RPC_TIMEOUT] (30s), and use it for the getblockheader
call in IsConfirmedBitcoinBlock and the getrawtransaction call in
CheckPeginSubsidyAndMinimum. This also bounds the advisory maturity check
in createrawpegin. The startup MainchainRPCCheck keeps the full
configured timeout. The help text now documents the cap.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant