github Galxe/gravity-sdk v1.6.0
Gravity v1.6.0

latest releases: v1.10.0, v1.9.3, v1.9.2...
5 months ago

Gravity v1.6.0

Release: v1.6.0

Full Changelog: gravity-testnet-v1.5.0...v1.6.0

Related pinned components:

This is the first mainnet release on the v1.6 line (gravity-mainnet-v1.6.0 tag aliases v1.6.0). The line forks from gravity-testnet-v1.5.0; the greth pin is intentionally rolled back one commit to drop the v1.5 Zeta-hardfork bytecode regeneration (gravity-reth#338) since Zeta is testnet-only.

Highlights

  • 50 Gwei minimum base fee (mainnet). The chainspec floor lands in gravity-reth#337 with a schedule-driven activation: config.gravityMinBaseFee in genesis carries the latest segment, historical segments live in branch-code so a binary always knows the full schedule. Production-pool admission and EIP-1559 next-block computation both clamp at the floor; non-Gravity chainspecs (e.g. Ethereum history sync) are untouched. Paired with SDK changes #689 / #692 / #693 to keep e2e fixtures and the CLI aligned.
  • Consensus crash-recovery hardening. Five separate panic / liveness fixes (#694, #697, #700, #701, #704) close the panic classes the audit and the staging-mainnet 2055785 stall surfaced. Validators now correctly recover from fast-forward-sync, epoch-change boundaries, mismatched aptos-db / reth heights, duplicate ordered-block pushes, and silently dropped commit votes.
  • Mempool broadcast cache rewrite. TxnCache no longer writes on the ingress path (#699) — duplicated/rejected txs no longer self-suppress their own re-arrival. The two-generation rotation is collapsed to a single TTL-bounded set with predictable memory.
  • Operator UX. New gravity_cli doctor self-diagnostic (#665) and per-node cluster knobs txpool_max_account_slots + vfn_discovery_method (#696) for shadow-fullnode topologies.
  • Release CI rebuilt (#710, #713) on the self-hosted Ubuntu 24.04 runner with native build, actionlint gate, and tag-trigger release-dryrun-<sha> artifacts.

What's Changed

Features

  • feat(cli): add doctor command for config and connectivity diagnostics by @ByteYue in #665
    • gravity-cli doctor runs a colored summary of config, rpc (chain_id + block number, 3s timeout), consensus-server (/dkg/status), deploy-path, version (CLI vs web3_clientVersion), and ports (8545 / 8551 / 9101). Flags: --output json, --skip-ports, --rpc-url / --server-url / --deploy-path overrides. Exits non-zero on any fail, so CI / pre-flight scripts can gate on it.
  • feat(cluster): support GCP Secret Manager identity for PFN nodes by @Lchangliang in #695
    • Adds GCP Secret Manager as an identity source for PFN cluster.toml deploys; staging/private mainnet PFNs no longer need keystore files on disk.
  • feat(cluster): per-node txpool slots and vfn-discovery overrides by @nekomoto911 in #696
    • txpool_max_account_slots (per-node, default 16) renders as --txpool.max-account-slots=N. Reth's upstream default 16 caps each sender's in-pool tx count for External-origin ingress; under Aptos shared-mempool propagation that throttles per-sender drain to ~16 mined tx/min steady-state — fine for e2e, painful for bench. Deployments needing the headroom opt in (txpool_max_account_slots = 10000).
    • vfn_discovery_method (per-node, falls back to discovery_method) controls only the validator's secondary full_node_networks/vfn block. Required when genesis registers shadow_fullnode — without it the vfn-network on-chain discovery resolves to the shadow VFN's identity and the validator's noise pubkey mismatch self-check loops every cycle.
  • feat(regression): pfn_chain stress suite + docker host-binary runtime image by @keanji-x in #709
    • New regression/pfn_chain_stress/ harness consolidates three earlier parallel bench dirs into one parameterized run.sh (--topology=chain|simple, --parallel=N, --cpuset). Dockerfile.host-binary wraps a pre-built host binary so the runtime image takes ~12 s to build vs ~30–65 min from source. Baselines on a 48-core / 256 GB testbed: ~5,800 chain TPS single cluster, ~14,500 simple ×3 parallel + cpuset.
  • feat: add genesis-alloc-audit agent skill by @nekomoto911 in #703
    • Adds an Agent skill that scans a Reth/geth-style genesis.json for dev/test credential leaks: Anvil/Hardhat default 10 accounts in alloc, dev chain IDs (1337 / 31337), literal dev mnemonic / privkey strings, and implausibly large "infinite money" balances. Runs as part of pre-mainnet ceremony checks.

Bug Fixes — Consensus

  • fix(consensus): prevent panic in VFN sync when parent blocks are pruned by @Lchangliang in #694
    • Two-part race fix. find_missing_randomness_block_on_path was starting sync from a stale/placeholder QC commit_info (epoch 0, round 0) when an unexecuted block was reached, landing on a pruned tree position; it now falls back to highest_commit_cert. append_blocks_for_sync no longer panics when concurrent tree pruning removes a parent block between fetch and insert — it warns and skips instead.
  • fix(consensus): use block_number limit in recover_blocks to prevent skipping epoch change by @Lchangliang in #697
    • The round-based epoch_change_limit_round check in recover_blocks would break before processing the only QC able to commit the epoch-change block, because that QC's commit_round exceeds the epoch-change round. Replaces with a block_number-based limit threaded into send_for_execution. Also fixes a "No LI found for root" panic on implicit-only commits: when reth lags aptos_db (e.g. crash between aptos commit and reth persist) the recovery root may have been committed implicitly via the 3-chain rule with no explicit commit LI; a placeholder WrappedLedgerInfo is constructed to pass BlockTree::new's assert and commit_callback overwrites it with the real LI during replay.
  • fix(consensus): prevent panic on empty ledger_infos in fast_forward_sync_by_epoch by @Lchangliang in #700
    • Early-return guards in fast_forward_sync_by_epoch: nothing to sync when all fetched blocks are ≤ local HCC round, and skip rebuild + epoch-change notification when ledger_infos is empty (the unwrap() on ledger_infos.last() was the direct cause of the sync_manager.rs:423 panic).
  • fix(consensus): bound commit_vote_cache by round window + reset validator pipeline after FFsync by @Richard1048576 in #701
    • Three-part architectural fix replacing the prior 1-line count-cap bump (1 000 → 100 000) that PR #645 introduced. Restructures the commit-vote cache from block_id → author → vote to Round → block_id → author → vote (drain leaks across forks fixed); bounds the cache by a configurable round window (highest_committed_round, highest_committed_round + max_pending_rounds) (default 100 ≈ 20 s cushion at staging cadence) and replies NACK for out-of-window votes so senders' rebroadcast_commit_votes_if_needed loop can retry; on validators, fully resets execution_client + reloads recovery_data + rebuilds blockstore after sync_to_highest_commit_cert fast-forwards to avoid the post-FFsync epoch-stuck symptom. Captured evidence on private mainnet: 761,475 dropped-vote WARNs combined across core-1/2/3 with the broken 1,000 cap. Also bumps gaptos 16ed314 → e9544c8c to pick up max_pending_rounds_in_commit_vote_cache config field.
  • fix(consensus): harden fast-forward recovery pipeline by @Lchangliang in #704
    • Direct follow-up to the staging-mainnet 2055785 stall. Aligns the consensus recovery root with execution when execution is at block 0 (genesis LI instead of a newer metadata-db LI). Only emits EpochChangeProof after the local commit root reaches the epoch-change boundary. Makes rand-manager ordered-block intake idempotent by tracking highest_dequeued_round and dropping stale ordered blocks pushed twice (covers the core-4 duplicate-send case where round 28843 blocked round 28844). Replays the ordered path after add_certs so a highest_ordered_cert already in BlockStore but not in the pipeline (the richard-1 case at rounds 28844..28851) is sent to execution. Broken execution-channel sends now surface as errors rather than silent Ok(()).
  • feat: Persist fetched quorum store batches to DB by @Lchangliang in #707
    • BatchReaderImpl::get_batch now calls a new save_fetched_batch_to_db after BatchRequester::request_batch returns, before the existing BatchStore::persist flow. Live cache/sign semantics are unchanged, but fullnodes that fetched a batch over RPC will keep a DB copy. Closes a recovery dead-end observed in production where validators had the QS payload but both VFNs were missing it locally, and the RPC node — which only requests batches from its VFN peers — could not recover.

Bug Fixes — Mempool

  • fix(mempool): drop add_txn cache write; simplify TxnCache to single generation by @keanji-x in #699
    • add_txn was inserting the tx hash into TxnCache before calling pool.add_external_txn, regardless of whether the reth pool accepted it. Reth rejections (nonce gap, balance, replacement failure, pool full) still occupied the dedup cache for 1×–2× TTL (default 60–120 s), so the same tx re-arriving via gossip during that window was silently dropped from re-broadcast — the node relied entirely on other peers to propagate. Fixed: only read_timeline writes the cache, on the drain path actually about to broadcast.
    • Also collapses the old_cache + cache two-generation rotation to a single set. The previous design's worst case was protecting "just-inserted" hashes from immediate eviction at the rotation boundary; the only consequence of losing that protection is one duplicate broadcast that gossip already dedupes. Trade-off removed: per-hash lifetime is now even (always TTL), peak memory drops to ~50% of before (single set), is_contains queries one set instead of two under the shared mutex.
  • fix(mempool): demote AlreadyImported to info, defer to reth's is_bad_transaction by @nekomoto911 in #705
    • bin/gravity_node/src/mempool.rs::add_external_txn was logging every reth PoolError at ERROR regardless of severity. On a chain stall, clients retry the same signed tx and aptos-mempool gossip re-broadcasts the duplicate — on the broadcast-hub validator (picked as BroadcastPeerPriority::Primary by multiple senders) this produced sustained ~1,100 ERROR/sec of PoolError { kind: AlreadyImported }. On the private-mainnet 2055785 stall, core-4 wrote 6.5 GB to debug.log in 5h41m while every other validator stayed under 70 KB. Fix: use reth's existing PoolError::is_bad_transaction() classifier; AlreadyImported, ReplacementUnderpriced, FeeCapBelowMinimumProtocolFeeCap, SpammerExceededCapacity, DiscardedOnInsert, ExistingConflictingTransactionType, Other are logged at INFO with "tx not added (recoverable)"; real pool-failure cases (e.g. bad signature) still fire at ERROR.

Bug Fixes — Cluster / E2E

  • fix(cluster): default genesisTimestampSecs so e2e survives genesis-tool strict mode by @ByteYue in #706
    • Pairs with contracts#90, which started requiring genesisTimestampSecs to actually set block.timestamp at genesis. Cluster scripts now default the field so e2e doesn't regress when the contracts repo's strict mode lands.

CI & Release Engineering

  • ci(release): runner-native release workflow + actionlint by @nekomoto911 in #710
    • New .github/workflows/release.yml building gravity_node / gravity_cli natively on the self-hosted Ubuntu 24.04 runner (no container). Trigger: push: tags: ['gravity-mainnet-*', 'gravity-testnet-*'] plus pull_request: paths: ['.github/workflows/release.yml'] for dry runs. Strategy: align the production-VM Packer image with the runner OS (Ubuntu 24.04, glibc 2.39), not the reverse — keeps the workflow simple (no per-run apt-install, target/ cache reuse) at the cost of requiring runner state stays in sync. Show build environment step logs /etc/os-release + ldd --version + toolchain versions every run so drift surfaces in the run record. Assert rustc matches rust-toolchain.toml fails fast on toolchain drift. Tag-push path requires the GH release to already exist (created via UI with human-authored notes); workflow only attaches binaries via gh release upload --clobber. PR-trigger path uploads as 7-day release-dryrun-<sha> artifact.
    • Adds .github/workflows/lint-workflows.yml running actionlint@1.7.7 on PRs touching .github/workflows/** (shellcheck temporarily disabled — pre-existing SC2086 / SC2053 noise in gate.yml / rust-ci.yml is out of scope).
  • ci(release): enable gcp-secret-manager feature for gravity_node by @Richard1048576 in #713
    • Stop-gap fix for #695's optional feature flag: v1.6.0 shipped with default = [] and the Staging Mainnet VFNs auto-restart-looped on aptos-config/.../network_config.rs:281 because their identity source is gcp_secret. The workflow now passes --features gcp-secret-manager for the gravity_node build step. (Superseded by #714 in v1.6.1, which makes the feature always-on at the dep level so the workflow can collapse back to a single build step.)

Chores

  • chore: gitignore Claude Code harness local state by @keanji-x in #698

gravity-reth changes since gravity-testnet-v1.5.0

The greth pin moved from fd3b0d64 to 364b8516. 364b8516 is actually an ancestor of fd3b0d64 — the v1.6.0 pin rolls back one greth commit (gravity-reth#338, regenerate Zeta StakePool bytecode for PR #73), since Zeta is a testnet-only hardfork artifact and v1.6.0 is the mainnet line baseline.

The greth commit that defines v1.6.0 is therefore #337 feat(fee): chainspec floor with code-driven activation schedule:

  • New ChainSpec.gravity_min_base_fee: Option<u64> parsed from genesis config.gravityMinBaseFee and ChainSpec.gravity_min_base_fee_activation_block: u64 from extra_fields. Presence of gravityMinBaseFee marks the chainspec as Gravity; non-Gravity chainspecs omit it and keep upstream EIP-1559 semantics.
  • New EthChainSpec::gravity_min_base_fee_at_block(block) -> Option<u64> trait method; default returns None (safe for non-Gravity / Ethereum history sync). ChainSpec impl encodes a single-segment [activation_block, ∞) schedule.
  • next_block_base_fee clamps the EIP-1559 result to the floor when the schedule is active for the next block.
  • EthereumPoolBuilder::build_pool raises the pool admission floor with .max(...) when the schedule is active for head + 1. Static decision at startup; admission missing tighten across activation only causes sub-floor txs to sit unincluded (pipe-exec rejects them at production), no consensus impact.
  • Reverts the broken parts of an earlier global-hardcoded 50 Gwei attempt (commit 7d0483e565) that broke validation of the real Ethereum mainnet London transition at block 12,965,000 where base_fee = 1 Gwei — validate_header_against_parent and the London-fork-boundary basefee in next_evm_env go back to upstream INITIAL_BASE_FEE.
  • MIN_PROTOCOL_BASE_FEE reverts to 7 wei so the upstream pool default works again; test-utility tx-gen / dev e2e tweaks are reverted.
  • Behaviour matrix: Ethereum mainnet history sync — no clamp; Gravity main — floor at 50 Gwei from block 0; released testnet branches — override activation height per network, pre-activation behaves like v1.4.

Contracts (gravity_chain_core_contracts) — window 2026-04-28 ... 2026-05-14

Don't miss a new gravity-sdk release

NewReleases is sending notifications on new releases.