Gravity v1.6.0
Release: v1.6.0
Full Changelog: gravity-testnet-v1.5.0...v1.6.0
Related pinned components:
- gravity-reth: rev
364b8516(feat(fee): chainspec floor with code-driven activation schedule, gravity-reth#337) - gravity_chain_core_contracts (window 2026-04-28 ... 2026-05-14): two genesis-tool fixes — #90
wire genesisTimestampSecs into EVM block.timestamp, #91make genesis JSON output byte-deterministic
This is the first mainnet release on the v1.6 line (gravity-mainnet-v1.6.0 tag aliases v1.6.0). The line forks from gravity-testnet-v1.5.0; the greth pin is intentionally rolled back one commit to drop the v1.5 Zeta-hardfork bytecode regeneration (gravity-reth#338) since Zeta is testnet-only.
Highlights
- 50 Gwei minimum base fee (mainnet). The chainspec floor lands in gravity-reth#337 with a schedule-driven activation:
config.gravityMinBaseFeein genesis carries the latest segment, historical segments live in branch-code so a binary always knows the full schedule. Production-pool admission and EIP-1559 next-block computation both clamp at the floor; non-Gravity chainspecs (e.g. Ethereum history sync) are untouched. Paired with SDK changes #689 / #692 / #693 to keep e2e fixtures and the CLI aligned. - Consensus crash-recovery hardening. Five separate panic / liveness fixes (
#694, #697, #700, #701, #704) close the panic classes the audit and the staging-mainnet 2055785 stall surfaced. Validators now correctly recover from fast-forward-sync, epoch-change boundaries, mismatched aptos-db / reth heights, duplicate ordered-block pushes, and silently dropped commit votes. - Mempool broadcast cache rewrite.
TxnCacheno longer writes on the ingress path (#699) — duplicated/rejected txs no longer self-suppress their own re-arrival. The two-generation rotation is collapsed to a single TTL-bounded set with predictable memory. - Operator UX. New
gravity_cli doctorself-diagnostic (#665) and per-node cluster knobstxpool_max_account_slots+vfn_discovery_method(#696) for shadow-fullnode topologies. - Release CI rebuilt (#710, #713) on the self-hosted Ubuntu 24.04 runner with native build, actionlint gate, and tag-trigger
release-dryrun-<sha>artifacts.
What's Changed
Features
- feat(cli): add doctor command for config and connectivity diagnostics by @ByteYue in #665
gravity-cli doctorruns a colored summary ofconfig,rpc(chain_id + block number, 3s timeout),consensus-server(/dkg/status),deploy-path,version(CLI vsweb3_clientVersion), andports(8545 / 8551 / 9101). Flags:--output json,--skip-ports,--rpc-url/--server-url/--deploy-pathoverrides. Exits non-zero on any fail, so CI / pre-flight scripts can gate on it.
- feat(cluster): support GCP Secret Manager identity for PFN nodes by @Lchangliang in #695
- Adds GCP Secret Manager as an identity source for PFN cluster.toml deploys; staging/private mainnet PFNs no longer need keystore files on disk.
- feat(cluster): per-node txpool slots and vfn-discovery overrides by @nekomoto911 in #696
txpool_max_account_slots(per-node, default 16) renders as--txpool.max-account-slots=N. Reth's upstream default 16 caps each sender's in-pool tx count forExternal-origin ingress; under Aptos shared-mempool propagation that throttles per-sender drain to ~16 mined tx/min steady-state — fine for e2e, painful for bench. Deployments needing the headroom opt in (txpool_max_account_slots = 10000).vfn_discovery_method(per-node, falls back todiscovery_method) controls only the validator's secondaryfull_node_networks/vfnblock. Required when genesis registersshadow_fullnode— without it the vfn-network on-chain discovery resolves to the shadow VFN's identity and the validator's noise pubkey mismatch self-check loops every cycle.
- feat(regression): pfn_chain stress suite + docker host-binary runtime image by @keanji-x in #709
- New
regression/pfn_chain_stress/harness consolidates three earlier parallel bench dirs into one parameterizedrun.sh(--topology=chain|simple,--parallel=N,--cpuset).Dockerfile.host-binarywraps a pre-built host binary so the runtime image takes ~12 s to build vs ~30–65 min from source. Baselines on a 48-core / 256 GB testbed: ~5,800 chain TPS single cluster, ~14,500 simple ×3 parallel + cpuset.
- New
- feat: add genesis-alloc-audit agent skill by @nekomoto911 in #703
- Adds an Agent skill that scans a Reth/geth-style
genesis.jsonfor dev/test credential leaks: Anvil/Hardhat default 10 accounts inalloc, dev chain IDs (1337 / 31337), literal dev mnemonic / privkey strings, and implausibly large "infinite money" balances. Runs as part of pre-mainnet ceremony checks.
- Adds an Agent skill that scans a Reth/geth-style
Bug Fixes — Consensus
- fix(consensus): prevent panic in VFN sync when parent blocks are pruned by @Lchangliang in #694
- Two-part race fix.
find_missing_randomness_block_on_pathwas starting sync from a stale/placeholder QCcommit_info(epoch 0, round 0) when an unexecuted block was reached, landing on a pruned tree position; it now falls back tohighest_commit_cert.append_blocks_for_syncno longer panics when concurrent tree pruning removes a parent block between fetch and insert — it warns and skips instead.
- Two-part race fix.
- fix(consensus): use block_number limit in recover_blocks to prevent skipping epoch change by @Lchangliang in #697
- The round-based
epoch_change_limit_roundcheck inrecover_blockswould break before processing the only QC able to commit the epoch-change block, because that QC'scommit_roundexceeds the epoch-change round. Replaces with ablock_number-based limit threaded intosend_for_execution. Also fixes a"No LI found for root"panic on implicit-only commits: when reth lags aptos_db (e.g. crash between aptos commit and reth persist) the recovery root may have been committed implicitly via the 3-chain rule with no explicit commit LI; a placeholderWrappedLedgerInfois constructed to passBlockTree::new's assert andcommit_callbackoverwrites it with the real LI during replay.
- The round-based
- fix(consensus): prevent panic on empty ledger_infos in fast_forward_sync_by_epoch by @Lchangliang in #700
- Early-return guards in
fast_forward_sync_by_epoch: nothing to sync when all fetched blocks are ≤ local HCC round, and skip rebuild + epoch-change notification whenledger_infosis empty (theunwrap()onledger_infos.last()was the direct cause of thesync_manager.rs:423panic).
- Early-return guards in
- fix(consensus): bound commit_vote_cache by round window + reset validator pipeline after FFsync by @Richard1048576 in #701
- Three-part architectural fix replacing the prior 1-line count-cap bump (
1 000 → 100 000) that PR #645 introduced. Restructures the commit-vote cache fromblock_id → author → votetoRound → block_id → author → vote(drain leaks across forks fixed); bounds the cache by a configurable round window(highest_committed_round, highest_committed_round + max_pending_rounds)(default 100 ≈ 20 s cushion at staging cadence) and replies NACK for out-of-window votes so senders'rebroadcast_commit_votes_if_neededloop can retry; on validators, fully resetsexecution_client+ reloadsrecovery_data+ rebuilds blockstore aftersync_to_highest_commit_certfast-forwards to avoid the post-FFsync epoch-stuck symptom. Captured evidence on private mainnet: 761,475 dropped-vote WARNs combined across core-1/2/3 with the broken 1,000 cap. Also bumps gaptos16ed314 → e9544c8cto pick upmax_pending_rounds_in_commit_vote_cacheconfig field.
- Three-part architectural fix replacing the prior 1-line count-cap bump (
- fix(consensus): harden fast-forward recovery pipeline by @Lchangliang in #704
- Direct follow-up to the staging-mainnet 2055785 stall. Aligns the consensus recovery root with execution when execution is at block 0 (genesis LI instead of a newer metadata-db LI). Only emits
EpochChangeProofafter the local commit root reaches the epoch-change boundary. Makes rand-manager ordered-block intake idempotent by trackinghighest_dequeued_roundand dropping stale ordered blocks pushed twice (covers the core-4 duplicate-send case where round 28843 blocked round 28844). Replays the ordered path afteradd_certsso ahighest_ordered_certalready in BlockStore but not in the pipeline (the richard-1 case at rounds 28844..28851) is sent to execution. Broken execution-channel sends now surface as errors rather than silentOk(()).
- Direct follow-up to the staging-mainnet 2055785 stall. Aligns the consensus recovery root with execution when execution is at block 0 (genesis LI instead of a newer metadata-db LI). Only emits
- feat: Persist fetched quorum store batches to DB by @Lchangliang in #707
BatchReaderImpl::get_batchnow calls a newsave_fetched_batch_to_dbafterBatchRequester::request_batchreturns, before the existingBatchStore::persistflow. Live cache/sign semantics are unchanged, but fullnodes that fetched a batch over RPC will keep a DB copy. Closes a recovery dead-end observed in production where validators had the QS payload but both VFNs were missing it locally, and the RPC node — which only requests batches from its VFN peers — could not recover.
Bug Fixes — Mempool
- fix(mempool): drop add_txn cache write; simplify TxnCache to single generation by @keanji-x in #699
add_txnwas inserting the tx hash intoTxnCachebefore callingpool.add_external_txn, regardless of whether the reth pool accepted it. Reth rejections (nonce gap, balance, replacement failure, pool full) still occupied the dedup cache for 1×–2× TTL (default 60–120 s), so the same tx re-arriving via gossip during that window was silently dropped from re-broadcast — the node relied entirely on other peers to propagate. Fixed: onlyread_timelinewrites the cache, on the drain path actually about to broadcast.- Also collapses the
old_cache+cachetwo-generation rotation to a single set. The previous design's worst case was protecting "just-inserted" hashes from immediate eviction at the rotation boundary; the only consequence of losing that protection is one duplicate broadcast that gossip already dedupes. Trade-off removed: per-hash lifetime is now even (always TTL), peak memory drops to ~50% of before (single set),is_containsqueries one set instead of two under the shared mutex.
- fix(mempool): demote AlreadyImported to info, defer to reth's is_bad_transaction by @nekomoto911 in #705
bin/gravity_node/src/mempool.rs::add_external_txnwas logging every rethPoolErrorat ERROR regardless of severity. On a chain stall, clients retry the same signed tx and aptos-mempool gossip re-broadcasts the duplicate — on the broadcast-hub validator (picked asBroadcastPeerPriority::Primaryby multiple senders) this produced sustained ~1,100 ERROR/sec ofPoolError { kind: AlreadyImported }. On the private-mainnet 2055785 stall, core-4 wrote 6.5 GB todebug.login 5h41m while every other validator stayed under 70 KB. Fix: use reth's existingPoolError::is_bad_transaction()classifier;AlreadyImported,ReplacementUnderpriced,FeeCapBelowMinimumProtocolFeeCap,SpammerExceededCapacity,DiscardedOnInsert,ExistingConflictingTransactionType,Otherare logged at INFO with"tx not added (recoverable)"; real pool-failure cases (e.g. bad signature) still fire at ERROR.
Bug Fixes — Cluster / E2E
- fix(cluster): default genesisTimestampSecs so e2e survives genesis-tool strict mode by @ByteYue in #706
- Pairs with contracts#90, which started requiring
genesisTimestampSecsto actually setblock.timestampat genesis. Cluster scripts now default the field so e2e doesn't regress when the contracts repo's strict mode lands.
- Pairs with contracts#90, which started requiring
CI & Release Engineering
- ci(release): runner-native release workflow + actionlint by @nekomoto911 in #710
- New
.github/workflows/release.ymlbuildinggravity_node/gravity_clinatively on the self-hosted Ubuntu 24.04 runner (no container). Trigger:push: tags: ['gravity-mainnet-*', 'gravity-testnet-*']pluspull_request: paths: ['.github/workflows/release.yml']for dry runs. Strategy: align the production-VM Packer image with the runner OS (Ubuntu 24.04, glibc 2.39), not the reverse — keeps the workflow simple (no per-run apt-install,target/cache reuse) at the cost of requiring runner state stays in sync.Show build environmentstep logs/etc/os-release+ldd --version+ toolchain versions every run so drift surfaces in the run record.Assert rustc matches rust-toolchain.tomlfails fast on toolchain drift. Tag-push path requires the GH release to already exist (created via UI with human-authored notes); workflow only attaches binaries viagh release upload --clobber. PR-trigger path uploads as 7-dayrelease-dryrun-<sha>artifact. - Adds
.github/workflows/lint-workflows.ymlrunningactionlint@1.7.7on PRs touching.github/workflows/**(shellcheck temporarily disabled — pre-existing SC2086 / SC2053 noise ingate.yml/rust-ci.ymlis out of scope).
- New
- ci(release): enable gcp-secret-manager feature for gravity_node by @Richard1048576 in #713
- Stop-gap fix for #695's optional feature flag: v1.6.0 shipped with
default = []and the Staging Mainnet VFNs auto-restart-looped onaptos-config/.../network_config.rs:281because their identity source isgcp_secret. The workflow now passes--features gcp-secret-managerfor thegravity_nodebuild step. (Superseded by #714 in v1.6.1, which makes the feature always-on at the dep level so the workflow can collapse back to a single build step.)
- Stop-gap fix for #695's optional feature flag: v1.6.0 shipped with
Chores
gravity-reth changes since gravity-testnet-v1.5.0
The greth pin moved from fd3b0d64 to 364b8516. 364b8516 is actually an ancestor of fd3b0d64 — the v1.6.0 pin rolls back one greth commit (gravity-reth#338, regenerate Zeta StakePool bytecode for PR #73), since Zeta is a testnet-only hardfork artifact and v1.6.0 is the mainnet line baseline.
The greth commit that defines v1.6.0 is therefore #337 feat(fee): chainspec floor with code-driven activation schedule:
- New
ChainSpec.gravity_min_base_fee: Option<u64>parsed from genesisconfig.gravityMinBaseFeeandChainSpec.gravity_min_base_fee_activation_block: u64fromextra_fields. Presence ofgravityMinBaseFeemarks the chainspec as Gravity; non-Gravity chainspecs omit it and keep upstream EIP-1559 semantics. - New
EthChainSpec::gravity_min_base_fee_at_block(block) -> Option<u64>trait method; default returnsNone(safe for non-Gravity / Ethereum history sync).ChainSpecimpl encodes a single-segment[activation_block, ∞)schedule. next_block_base_feeclamps the EIP-1559 result to the floor when the schedule is active for the next block.EthereumPoolBuilder::build_poolraises the pool admission floor with.max(...)when the schedule is active forhead + 1. Static decision at startup; admission missing tighten across activation only causes sub-floor txs to sit unincluded (pipe-exec rejects them at production), no consensus impact.- Reverts the broken parts of an earlier global-hardcoded 50 Gwei attempt (commit
7d0483e565) that broke validation of the real Ethereum mainnet London transition at block 12,965,000 wherebase_fee = 1 Gwei—validate_header_against_parentand the London-fork-boundary basefee innext_evm_envgo back to upstreamINITIAL_BASE_FEE. MIN_PROTOCOL_BASE_FEEreverts to 7 wei so the upstream pool default works again; test-utility tx-gen / dev e2e tweaks are reverted.- Behaviour matrix: Ethereum mainnet history sync — no clamp; Gravity main — floor at 50 Gwei from block 0; released testnet branches — override activation height per network, pre-activation behaves like v1.4.
Contracts (gravity_chain_core_contracts) — window 2026-04-28 ... 2026-05-14
#90 fix(genesis-tool): wire genesisTimestampSecs into EVM block.timestamp— closes the divergence between Aptos block-0 timestamp and EVMblock.timestampthat previously left state-dependent system contracts (StakingConfig.lockupDurationMicros, etc.) reading 0 at genesis.#91 fix(genesis-tool): make genesis JSON output byte-deterministic— required for mainnet ceremony: identical inputs must produce a byte-identicalgenesis.jsonacross signing parties.