mirror of
https://github.com/tendermint/tendermint.git
synced 2026-10-01 04:05:45 +00:00
docs: migrate adrs and rfcs across from master (#9115)
This commit is contained in:
@@ -0,0 +1,63 @@
|
||||
---
|
||||
order: 1
|
||||
parent:
|
||||
order: false
|
||||
---
|
||||
|
||||
# Requests for Comments
|
||||
|
||||
A Request for Comments (RFC) is a record of discussion on an open-ended topic
|
||||
related to the design and implementation of Tendermint Core, for which no
|
||||
immediate decision is required.
|
||||
|
||||
The purpose of an RFC is to serve as a historical record of a high-level
|
||||
discussion that might otherwise only be recorded in an ad hoc way (for example,
|
||||
via gists or Google docs) that are difficult to discover for someone after the
|
||||
fact. An RFC _may_ give rise to more specific architectural _decisions_ for
|
||||
Tendermint, but those decisions must be recorded separately in [Architecture
|
||||
Decision Records (ADR)](./../architecture).
|
||||
|
||||
As a rule of thumb, if you can articulate a specific question that needs to be
|
||||
answered, write an ADR. If you need to explore the topic and get input from
|
||||
others to know what questions need to be answered, an RFC may be appropriate.
|
||||
|
||||
## RFC Content
|
||||
|
||||
An RFC should provide:
|
||||
|
||||
- A **changelog**, documenting when and how the RFC has changed.
|
||||
- An **abstract**, briefly summarizing the topic so the reader can quickly tell
|
||||
whether it is relevant to their interest.
|
||||
- Any **background** a reader will need to understand and participate in the
|
||||
substance of the discussion (links to other documents are fine here).
|
||||
- The **discussion**, the primary content of the document.
|
||||
|
||||
The [rfc-template.md](./rfc-template.md) file includes placeholders for these
|
||||
sections.
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [RFC-000: P2P Roadmap](./rfc-000-p2p-roadmap.rst)
|
||||
- [RFC-001: Storage Engines](./rfc-001-storage-engine.rst)
|
||||
- [RFC-002: Interprocess Communication](./rfc-002-ipc-ecosystem.md)
|
||||
- [RFC-003: Performance Taxonomy](./rfc-003-performance-questions.md)
|
||||
- [RFC-004: E2E Test Framework Enhancements](./rfc-004-e2e-framework.rst)
|
||||
- [RFC-005: Event System](./rfc-005-event-system.rst)
|
||||
- [RFC-006: Event Subscription](./rfc-006-event-subscription.md)
|
||||
- [RFC-007: Deterministic Proto Byte Serialization](./rfc-007-deterministic-proto-bytes.md)
|
||||
- [RFC-008: Don't Panic](./rfc-008-do-not-panic.md)
|
||||
- [RFC-009: Consensus Parameter Upgrades](./rfc-009-consensus-parameter-upgrades.md)
|
||||
- [RFC-010: P2P Light Client](./rfc-010-p2p-light-client.rst)
|
||||
- [RFC-011: Delete Gas](./rfc-011-delete-gas.md)
|
||||
- [RFC-012: Event Indexing Revisited](./rfc-012-custom-indexing.md)
|
||||
- [RFC-013: ABCI++](./rfc-013-abci++.md)
|
||||
- [RFC-014: Semantic Versioning](./rfc-014-semantic-versioning.md)
|
||||
- [RFC-015: ABCI++ Tx Mutation](./rfc-015-abci++-tx-mutation.md)
|
||||
- [RFC-016: Node Architecture](./rfc-016-node-architecture.md)
|
||||
- [RFC-017: ABCI++ Vote Extension Propagation](./rfc-017-abci++-vote-extension-propag.md)
|
||||
- [RFC-018: BLS Signature Aggregation Exploration](./rfc-018-bls-agg-exploration.md)
|
||||
- [RFC-019: Configuration File Versioning](./rfc-019-config-version.md)
|
||||
- [RFC-020: Onboarding Projects](./rfc-020-onboarding-projects.rst)
|
||||
- [RFC-021: The Future of the Socket Protocol](./rfc-021-socket-protocol.md)
|
||||
|
||||
<!-- - [RFC-NNN: Title](./rfc-NNN-title.md) -->
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 2.7 MiB |
Binary file not shown.
|
After Width: | Height: | Size: 2.5 MiB |
File diff suppressed because one or more lines are too long
|
After Width: | Height: | Size: 24 KiB |
@@ -0,0 +1,316 @@
|
||||
====================
|
||||
RFC 000: P2P Roadmap
|
||||
====================
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2021-08-20: Completed initial draft and distributed via a gist
|
||||
- 2021-08-25: Migrated as an RFC and changed format
|
||||
|
||||
Abstract
|
||||
--------
|
||||
|
||||
This document discusses the future of peer network management in Tendermint, with
|
||||
a particular focus on features, semantics, and a proposed roadmap.
|
||||
Specifically, we consider libp2p as a tool kit for implementing some fundamentals.
|
||||
|
||||
Background
|
||||
----------
|
||||
|
||||
For the 0.35 release cycle the switching/routing layer of Tendermint was
|
||||
replaced. This work was done "in place," and produced a version of Tendermint
|
||||
that was backward-compatible and interoperable with previous versions of the
|
||||
software. While there are new p2p/peer management constructs in the new
|
||||
version (e.g. ``PeerManager`` and ``Router``), the main effect of this change
|
||||
was to simplify the ways that other components within Tendermint interacted with
|
||||
the peer management layer, and to make it possible for higher-level components
|
||||
(specifically the reactors), to be used and tested more independently.
|
||||
|
||||
This refactoring, which was a major undertaking, was entirely necessary to
|
||||
enable areas for future development and iteration on this aspect of
|
||||
Tendermint. There are also a number of potential user-facing features that
|
||||
depend heavily on the p2p layer: additional transport protocols, transport
|
||||
compression, improved resilience to network partitions. These improvements to
|
||||
modularity, stability, and reliability of the p2p system will also make
|
||||
ongoing maintenance and feature development easier in the rest of Tendermint.
|
||||
|
||||
Critique of Current Peer-to-Peer Infrastructure
|
||||
---------------------------------------
|
||||
|
||||
The current (refactored) P2P stack is an improvement on the previous iteration
|
||||
(legacy), but as of 0.35, there remains room for improvement in the design and
|
||||
implementation of the P2P layer.
|
||||
|
||||
Some limitations of the current stack include:
|
||||
|
||||
- heavy reliance on buffering to avoid backups in the flow of components,
|
||||
which is fragile to maintain and can lead to unexpected memory usage
|
||||
patterns and forces the routing layer to make decisions about when messages
|
||||
should be discarded.
|
||||
|
||||
- the current p2p stack relies on convention (rather than the compiler) to
|
||||
enforce the API boundaries and conventions between reactors and the router,
|
||||
making it very easy to write "wrong" reactor code or introduce a bad
|
||||
dependency.
|
||||
|
||||
- the current stack is probably more complex and difficult to maintain because
|
||||
the legacy system must coexist with the new components in 0.35. When the
|
||||
legacy stack is removed there are some simple changes that will become
|
||||
possible and could reduce the complexity of the new system. (e.g. `#6598
|
||||
<https://github.com/tendermint/tendermint/issues/6598>`_.)
|
||||
|
||||
- the current stack encapsulates a lot of information about peers, and makes it
|
||||
difficult to expose that information to monitoring/observability tools. This
|
||||
general opacity also makes it difficult to interact with the peer system
|
||||
from other areas of the code base (e.g. tests, reactors).
|
||||
|
||||
- the legacy stack provided some control to operators to force the system to
|
||||
dial new peers or seed nodes or manipulate the topology of the system _in
|
||||
situ_. The current stack can't easily provide this, and while the new stack
|
||||
may have better behavior, it does leave operators hands tied.
|
||||
|
||||
Some of these issues will be resolved early in the 0.36 cycle, with the
|
||||
removal of the legacy components.
|
||||
|
||||
The 0.36 release also provides the opportunity to make changes to the
|
||||
protocol, as the release will not be compatible with previous releases.
|
||||
|
||||
Areas for Development
|
||||
---------------------
|
||||
|
||||
These sections describe features that may make sense to include in a Phase 2 of
|
||||
a P2P project.
|
||||
|
||||
Internal Message Passing
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Currently, there's no provision for intranode communication using the P2P
|
||||
layer, which means when two reactors need to interact with each other they
|
||||
have to have dependencies on each other's interfaces, and
|
||||
initialization. Changing these interactions (e.g. transitions between
|
||||
blocksync and consensus) from procedure calls to message passing.
|
||||
|
||||
This is a relatively simple change and could be implemented with the following
|
||||
components:
|
||||
|
||||
- a constant to represent "local" delivery as the ``To`` field on
|
||||
``p2p.Envelope``.
|
||||
|
||||
- special path for routing local messages that doesn't require message
|
||||
serialization (protobuf marshalling/unmarshaling).
|
||||
|
||||
Adding these semantics, particularly if in conjunction with synchronous
|
||||
semantics provides a solution to dependency graph problems currently present
|
||||
in the Tendermint codebase, which will simplify development, make it possible
|
||||
to isolate components for testing.
|
||||
|
||||
Eventually, this will also make it possible to have a logical Tendermint node
|
||||
running in multiple processes or in a collection of containers, although the
|
||||
usecase of this may be debatable.
|
||||
|
||||
Synchronous Semantics (Paired Request/Response)
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
In the current system, all messages are sent with fire-and-forget semantics,
|
||||
and there's no coupling between a request sent via the p2p layer, and a
|
||||
response. These kinds of semantics would simplify the implementation of
|
||||
state and block sync reactors, and make intra-node message passing more
|
||||
powerful.
|
||||
|
||||
For some interactions, like gossiping transactions between the mempools of
|
||||
different nodes, fire-and-forget semantics make sense, but for other
|
||||
operations the missing link between requests/responses leads to either
|
||||
inefficiency when a node fails to respond or becomes unavailable, or code that
|
||||
is just difficult to follow.
|
||||
|
||||
To support this kind of work, the protocol would need to accommodate some kind
|
||||
of request/response ID to allow identifying out-of-order responses over a
|
||||
single connection. Additionally, expanded the programming model of the
|
||||
``p2p.Channel`` to accommodate some kind of _future_ or similar paradigm to
|
||||
make it viable to write reactor code without needing for the reactor developer
|
||||
to wrestle with lower level concurrency constructs.
|
||||
|
||||
|
||||
Timeout Handling (QoS)
|
||||
~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Currently, all timeouts, buffering, and QoS features are handled at the router
|
||||
layer, and the reactors are implemented in ways that assume/require
|
||||
asynchronous operation. This both increases the required complexity at the
|
||||
routing layer, and means that misbehavior at the reactor level is difficult to
|
||||
detect or attribute. Additionally, the current system provides three main
|
||||
parameters to control quality of service:
|
||||
|
||||
- buffer sizes for channels and queues.
|
||||
|
||||
- priorities for channels
|
||||
|
||||
- queue implementation details for shedding load.
|
||||
|
||||
These end up being quite coarse controls, and changing the settings are
|
||||
difficult because as the queues and channels are able to buffer large numbers
|
||||
of messages it can be hard to see the impact of a given change, particularly
|
||||
in our extant test environment. In general, we should endeavor to:
|
||||
|
||||
- set real timeouts, via contexts, on most message send operations, so that
|
||||
senders rather than queues can be responsible for timeout
|
||||
logic. Additionally, this will make it possible to avoid sending messages
|
||||
during shutdown.
|
||||
|
||||
- reduce (to the greatest extent possible) the amount of buffering in
|
||||
channels and the queues, to more readily surface backpressure and reduce the
|
||||
potential for buildup of stale messages.
|
||||
|
||||
Stream Based Connection Handling
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Currently the transport layer is message based, which makes sense from a
|
||||
mental model of how the protocol works, but makes it more difficult to
|
||||
implement transports and connection types, as it forces a higher level view of
|
||||
the connection and interaction which makes it harder to implement for novel
|
||||
transport types and makes it more likely that message-based caching and rate
|
||||
limiting will be implemented at the transport layer rather than at a more
|
||||
appropriate level.
|
||||
|
||||
The transport then, would be responsible for negotiating the connection and the
|
||||
handshake and otherwise behave like a socket/file descriptor with ``Read`` and
|
||||
``Write`` methods.
|
||||
|
||||
While this was included in the initial design for the new P2P layer, it may be
|
||||
obviated entirely if the transport and peer layer is replaced with libp2p,
|
||||
which is primarily stream based.
|
||||
|
||||
Service Discovery
|
||||
~~~~~~~~~~~~~~~~~
|
||||
|
||||
In the current system, Tendermint assumes that all nodes in a network are
|
||||
largely equivalent, and nodes tend to be "chatty" making many requests of
|
||||
large numbers of peers and waiting for peers to (hopefully) respond. While
|
||||
this works and has allowed Tendermint to get to a certain point, this both
|
||||
produces a theoretical scaling bottle neck and makes it harder to test and
|
||||
verify components of the system.
|
||||
|
||||
In addition to peer's identity and connection information, peers should be
|
||||
able to advertise a number of services or capabilities, and node operators or
|
||||
developers should be able to specify peer capability requirements (e.g. target
|
||||
at least <x>-percent of peers with <y> capability.)
|
||||
|
||||
These capabilities may be useful in selecting peers to send messages to, it
|
||||
may make sense to extend Tendermint's message addressing capability to allow
|
||||
reactors to send messages to groups of peers based on role rather than only
|
||||
allowing addressing to one or all peers.
|
||||
|
||||
Having a good service discovery mechanism may pair well with the synchronous
|
||||
semantics (request/response) work, as it allows reactors to "make a request of
|
||||
a peer with <x> capability and wait for the response," rather force the
|
||||
reactors to need to track the capabilities or state of specific peers.
|
||||
|
||||
Solutions
|
||||
---------
|
||||
|
||||
Continued Homegrown Implementation
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
The current peer system is homegrown and is conceptually compatible with the
|
||||
needs of the project, and while there are limitations to the system, the p2p
|
||||
layer is not (currently as of 0.35) a major source of bugs or friction during
|
||||
development.
|
||||
|
||||
However, the current implementation makes a number of allowances for
|
||||
interoperability, and there are a collection of iterative improvements that
|
||||
should be considered in the next couple of releases. To maintain the current
|
||||
implementation, upcoming work would include:
|
||||
|
||||
- change the ``Transport`` mechanism to facilitate easier implementations.
|
||||
|
||||
- implement different ``Transport`` handlers to be able to manage peer
|
||||
connections using different protocols (e.g. QUIC, etc.)
|
||||
|
||||
- entirely remove the constructs and implementations of the legacy peer
|
||||
implementation.
|
||||
|
||||
- establish and enforce clearer chains of responsibility for connection
|
||||
establishment (e.g. handshaking, setup,) which is currently shared between
|
||||
three components.
|
||||
|
||||
- report better metrics regarding the into the state of peers and network
|
||||
connectivity, which are opaque outside of the system. This is constrained at
|
||||
the moment as a side effect of the split responsibility for connection
|
||||
establishment.
|
||||
|
||||
- extend the PEX system to include service information so that nodes in the
|
||||
network weren't necessarily homogeneous.
|
||||
|
||||
While maintaining a bespoke peer management layer would seem to distract from
|
||||
development of core functionality, the truth is that (once the legacy code is
|
||||
removed,) the scope of the peer layer is relatively small from a maintenance
|
||||
perspective, and having control at this layer might actually afford the
|
||||
project with the ability to more rapidly iterate on some features.
|
||||
|
||||
LibP2P
|
||||
~~~~~~
|
||||
|
||||
LibP2P provides components that, approximately, account for the
|
||||
``PeerManager`` and ``Transport`` components of the current (new) P2P
|
||||
stack. The Go APIs seem reasonable, and being able to externalize the
|
||||
implementation details of peer and connection management seems like it could
|
||||
provide a lot of benefits, particularly in supporting a more active ecosystem.
|
||||
|
||||
In general the API provides the kind of stream-based, multi-protocol
|
||||
supporting, and idiomatic baseline for implementing a peer layer. Additionally
|
||||
because it handles peer exchange and connection management at a lower
|
||||
level, by using libp2p it'd be possible to remove a good deal of code in favor
|
||||
of just using libp2p. Having said that, Tendermint's P2P layer covers a
|
||||
greater scope (e.g. message routing to different peers) and that layer is
|
||||
something that Tendermint might want to retain.
|
||||
|
||||
The are a number of unknowns that require more research including how much of
|
||||
a peer database the Tendermint engine itself needs to maintain, in order to
|
||||
support higher level operations (consensus, statesync), but it might be the
|
||||
case that our internal systems need to know much less about peers than
|
||||
otherwise specified. Similarly, the current system has a notion of peer
|
||||
scoring that cannot be communicated to libp2p, which may be fine as this is
|
||||
only used to support peer exchange (PEX,) which would become a property libp2p
|
||||
and not expressed in it's current higher-level form.
|
||||
|
||||
In general, the effort to switch to libp2p would involve:
|
||||
|
||||
- timing it during an appropriate protocol-breaking window, as it doesn't seem
|
||||
viable to support both libp2p *and* the current p2p protocol.
|
||||
|
||||
- providing some in-memory testing network to support the use case that the
|
||||
current ``p2p.MemoryNetwork`` provides.
|
||||
|
||||
- re-homing the ``p2p.Router`` implementation on top of libp2p components to
|
||||
be able to maintain the current reactor implementations.
|
||||
|
||||
Open question include:
|
||||
|
||||
- how much local buffering should we be doing? It sort of seems like we should
|
||||
figure out what the expected behavior is for libp2p for QoS-type
|
||||
functionality, and if our requirements mean that we should be implementing
|
||||
this on top of things ourselves?
|
||||
|
||||
- if Tendermint was going to use libp2p, how would libp2p's stability
|
||||
guarantees (protocol, etc.) impact/constrain Tendermint's stability
|
||||
guarantees?
|
||||
|
||||
- what kind of introspection does libp2p provide, and to what extend would
|
||||
this change or constrain the kind of observability that Tendermint is able
|
||||
to provide?
|
||||
|
||||
- how do efforts to select "the best" (healthy, close, well-behaving, etc.)
|
||||
peers work out if Tendermint is not maintaining a local peer database?
|
||||
|
||||
- would adding additional higher level semantics (internal message passing,
|
||||
request/response pairs, service discovery, etc.) facilitate removing some of
|
||||
the direct linkages between constructs/components in the system and reduce
|
||||
the need for Tendermint nodes to maintain state about its peers?
|
||||
|
||||
References
|
||||
----------
|
||||
|
||||
- `Tracking Ticket for P2P Refactor Project <https://github.com/tendermint/tendermint/issues/5670>`_
|
||||
- `ADR 61: P2P Refactor Scope <../architecture/adr-061-p2p-refactor-scope.md>`_
|
||||
- `ADR 62: P2P Architecture and Abstraction <../architecture/adr-061-p2p-architecture.md>`_
|
||||
@@ -0,0 +1,179 @@
|
||||
===========================================
|
||||
RFC 001: Storage Engines and Database Layer
|
||||
===========================================
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2021-04-19: Initial Draft (gist)
|
||||
- 2021-09-02: Migrated to RFC folder, with some updates
|
||||
|
||||
Abstract
|
||||
--------
|
||||
|
||||
The aspect of Tendermint that's responsible for persistence and storage (often
|
||||
"the database" internally) represents a bottle neck in the architecture of the
|
||||
platform, that the 0.36 release presents a good opportunity to correct. The
|
||||
current storage engine layer provides a great deal of flexibility that is
|
||||
difficult for users to leverage or benefit from, while also making it harder
|
||||
for Tendermint Core developers to deliver improvements on storage engine. This
|
||||
RFC discusses the possible improvements to this layer of the system.
|
||||
|
||||
Background
|
||||
----------
|
||||
|
||||
Tendermint has a very thin common wrapper that makes Tendermint itself
|
||||
(largely) agnostic to the data storage layer (within the realm of the popular
|
||||
key-value/embedded databases.) This flexibility is not particularly useful:
|
||||
the benefits of a specific database engine in the context of Tendermint is not
|
||||
particularly well understood, and the maintenance burden for multiple backends
|
||||
is not commensurate with the benefit provided. Additionally, because the data
|
||||
storage layer is handled generically, and most tests run with an in-memory
|
||||
framework, it's difficult to take advantage of any higher-level features of a
|
||||
database engine.
|
||||
|
||||
Ideally, developers within Tendermint will be able to interact with persisted
|
||||
data via an interface that can function, approximately like an object
|
||||
store, and this storage interface will be able to accommodate all existing
|
||||
persistence workloads (e.g. block storage, local peer management information
|
||||
like the "address book", crash-recovery log like the WAL.) In addition to
|
||||
providing a more ergonomic interface and new semantics, by selecting a single
|
||||
storage engine tendermint can use native durability and atomicity features of
|
||||
the storage engine and simplify its own implementations.
|
||||
|
||||
Data Access Patterns
|
||||
~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Tendermint's data access patterns have the following characteristics:
|
||||
|
||||
- aggregate data size often exceeds memory.
|
||||
|
||||
- data is rarely mutated after it's written for most data (e.g. blocks), but
|
||||
small amounts of working data is persisted by nodes and is frequently
|
||||
mutated (e.g. peer information, validator information.)
|
||||
|
||||
- read patterns can be quite random.
|
||||
|
||||
- crash resistance and crash recovery, provided by write-ahead-logs (in
|
||||
consensus, and potentially for the mempool) should allow the system to
|
||||
resume work after an unexpected shut down.
|
||||
|
||||
Project Goals
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
As we think about replacing the current persistence layer, we should consider
|
||||
the following high level goals:
|
||||
|
||||
- drop dependencies on storage engines that have a CGo dependency.
|
||||
|
||||
- encapsulate data format and data storage from higher-level services
|
||||
(e.g. reactors) within tendermint.
|
||||
|
||||
- select a storage engine that does not incur any additional operational
|
||||
complexity (e.g. database should be embedded.)
|
||||
|
||||
- provide database semantics with sufficient ACID, snapshots, and
|
||||
transactional support.
|
||||
|
||||
Open Questions
|
||||
~~~~~~~~~~~~~~
|
||||
|
||||
The following questions remain:
|
||||
|
||||
- what kind of data-access concurrency does tendermint require?
|
||||
|
||||
- would tendermint users SDK/etc. benefit from some shared database
|
||||
infrastructure?
|
||||
|
||||
- In earlier conversations it seemed as if the SDK has selected Badger and
|
||||
RocksDB for their storage engines, and it might make sense to be able to
|
||||
(optionally) pass a handle to a Badger instance between the libraries in
|
||||
some cases.
|
||||
|
||||
- what are typical data sizes, and what kinds of memory sizes can we expect
|
||||
operators to be able to provide?
|
||||
|
||||
- in addition to simple persistence, what kind of additional semantics would
|
||||
tendermint like to enjoy (e.g. transactional semantics, unique constraints,
|
||||
indexes, in-place-updates, etc.)?
|
||||
|
||||
Decision Framework
|
||||
~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Given the constraint of removing the CGo dependency, the decision is between
|
||||
"badger" and "boltdb" (in the form of the etcd/CoreOS fork,) as low level. On
|
||||
top of this and somewhat orthogonally, we must also decide on the interface to
|
||||
the database and how the larger application will have to interact with the
|
||||
database layer. Users of the data layer shouldn't ever need to interact with
|
||||
raw byte slices from the database, and should mostly have the experience of
|
||||
interacting with Go-types.
|
||||
|
||||
Badger is more consistently developed and has a broader feature set than
|
||||
Bolt. At the same time, Badger is likely more memory intensive and may have
|
||||
more overhead in terms of open file handles given it's model. At first glance,
|
||||
Badger is the obvious choice: it's actively developed and it has a lot of
|
||||
features that could be useful. Bolt is not without some benefits: it's stable
|
||||
and is maintained by the etcd folks, it's simpler model (single memory mapped
|
||||
file, etc,) may be easier to reason about.
|
||||
|
||||
I propose that we consider the following specific questions about storage
|
||||
engines:
|
||||
|
||||
- does Badger's evolving development, which may result in data file format
|
||||
changes in the future, and could restrict our access to using the latest
|
||||
version of the library between major upgrades, present a problem?
|
||||
|
||||
- do we do we have goals/concerns about memory footprint that Badger may
|
||||
prevent us from hitting, particularly as data sets grow over time?
|
||||
|
||||
- what kind of additional tooling might we need/like to build (dump/restore,
|
||||
etc.)?
|
||||
|
||||
- do we want to run unit/integration tests against a data files on disk rather
|
||||
than relying exclusively on the memory database?
|
||||
|
||||
Project Scope
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
This project will consist of the following aspects:
|
||||
|
||||
- selecting a storage engine, and modifying the tendermint codebase to
|
||||
disallow any configuration of the storage engine outside of the tendermint.
|
||||
|
||||
- remove the dependency on the current tm-db interfaces and replace with some
|
||||
internalized, safe, and ergonomic interface for data persistence with all
|
||||
required database semantics.
|
||||
|
||||
- update core tendermint code to use the new interface and data tools.
|
||||
|
||||
Next Steps
|
||||
~~~~~~~~~~
|
||||
|
||||
- circulate the RFC, and discuss options with appropriate stakeholders.
|
||||
|
||||
- write brief ADR to summarize decisions around technical decisions reached
|
||||
during the RFC phase.
|
||||
|
||||
References
|
||||
----------
|
||||
|
||||
- `bolddb <https://github.com/etcd-io/bbolt>`_
|
||||
- `badger <https://github.com/dgraph-io/badger>`_
|
||||
- `badgerdb overview <https://dbdb.io/db/badgerdb>`_
|
||||
- `botldb overview <https://dbdb.io/db/boltdb>`_
|
||||
- `boltdb vs badger <https://tech.townsourced.com/post/boltdb-vs-badger>`_
|
||||
- `bolthold <https://github.com/timshannon/bolthold>`_
|
||||
- `badgerhold <https://github.com/timshannon/badgerhold>`_
|
||||
- `Pebble <https://github.com/cockroachdb/pebble>`_
|
||||
- `SDK Issue Regarding IVAL <https://github.com/cosmos/cosmos-sdk/issues/7100>`_
|
||||
- `SDK Discussion about SMT/IVAL <https://github.com/cosmos/cosmos-sdk/discussions/8297>`_
|
||||
|
||||
Discussion
|
||||
----------
|
||||
|
||||
- All things being equal, my tendency would be to use badger, with badgerhold
|
||||
(if that makes sense) for its ergonomics and indexing capabilities, which
|
||||
will require some small selection of wrappers for better write transaction
|
||||
support. This is a weakly held tendency/belief and I think it would be
|
||||
useful for the RFC process to build consensus (or not) around this basic
|
||||
assumption.
|
||||
@@ -0,0 +1,420 @@
|
||||
# RFC 002: Interprocess Communication (IPC) in Tendermint
|
||||
|
||||
## Changelog
|
||||
|
||||
- 08-Sep-2021: Initial draft (@creachadair).
|
||||
|
||||
|
||||
## Abstract
|
||||
|
||||
Communication in Tendermint among consensus nodes, applications, and operator
|
||||
tools all use different message formats and transport mechanisms. In some
|
||||
cases there are multiple options. Having all these options complicates both the
|
||||
code and the developer experience, and hides bugs. To support a more robust,
|
||||
trustworthy, and usable system, we should document which communication paths
|
||||
are essential, which could be removed or reduced in scope, and what we can
|
||||
improve for the most important use cases.
|
||||
|
||||
This document proposes a variety of possible improvements of varying size and
|
||||
scope. Specific design proposals should get their own documentation.
|
||||
|
||||
|
||||
## Background
|
||||
|
||||
The Tendermint state replication engine has a complex IPC footprint.
|
||||
|
||||
1. Consensus nodes communicate with each other using a networked peer-to-peer
|
||||
message-passing protocol.
|
||||
|
||||
2. Consensus nodes communicate with the application whose state is being
|
||||
replicated via the [Application BlockChain Interface (ABCI)][abci].
|
||||
|
||||
3. Consensus nodes export a network-accessible [RPC service][rpc-service] to
|
||||
support operations (bootstrapping, debugging) and synchronization of [light clients][light-client].
|
||||
This interface is also used by the [`tendermint` CLI][tm-cli].
|
||||
|
||||
4. Consensus nodes export a gRPC service exposing a subset of the methods of
|
||||
the RPC service described by (3). This was intended to simplify the
|
||||
implementation of tools that already use gRPC to communicate with an
|
||||
application (via the Cosmos SDK), and wanted to also talk to the consensus
|
||||
node without implementing yet another RPC protocol.
|
||||
|
||||
The gRPC interface to the consensus node has been deprecated and is slated
|
||||
for removal in the forthcoming Tendermint v0.36 release.
|
||||
|
||||
5. Consensus nodes may optionally communicate with a "remote signer" that holds
|
||||
a validator key and can provide public keys and signatures to the consensus
|
||||
node. One of the stated goals of this configuration is to allow the signer
|
||||
to be run on a private network, separate from the consensus node, so that a
|
||||
compromise of the consensus node from the public network would be less
|
||||
likely to expose validator keys.
|
||||
|
||||
## Discussion: Transport Mechanisms
|
||||
|
||||
### Remote Signer Transport
|
||||
|
||||
A remote signer communicates with the consensus node in one of two ways:
|
||||
|
||||
1. "Raw": Using a TCP or Unix-domain socket which carries varint-prefixed
|
||||
protocol buffer messages. In this mode, the consensus node is the server,
|
||||
and the remote signer is the client.
|
||||
|
||||
This mode has been deprecated, and is intended to be removed.
|
||||
|
||||
2. gRPC: This mode uses the same protobuf messages as "Raw" node, but uses a
|
||||
standard encrypted gRPC HTTP/2 stub as the transport. In this mode, the
|
||||
remote signer is the server and the consensus node is the client.
|
||||
|
||||
|
||||
### ABCI Transport
|
||||
|
||||
In ABCI, the _application_ is the server, and the Tendermint consensus engine
|
||||
is the client. Most applications implement the server using the [Cosmos SDK][cosmos-sdk],
|
||||
which handles low-level details of the ABCI interaction and provides a
|
||||
higher-level interface to the rest of the application. The SDK is written in Go.
|
||||
|
||||
Beneath the SDK, the application communicates with Tendermint core in one of
|
||||
two ways:
|
||||
|
||||
- In-process direct calls (for applications written in Go and compiled against
|
||||
the Tendermint code). This is an optimization for the common case where an
|
||||
application is written in Go, to save on the overhead of marshaling and
|
||||
unmarshaling requests and responses within the same process:
|
||||
[`abci/client/local_client.go`][local-client]
|
||||
|
||||
- A custom remote procedure protocol built on wire-format protobuf messages
|
||||
using a socket (the "socket protocol"): [`abci/server/socket_server.go`][socket-server]
|
||||
|
||||
The SDK also provides a [gRPC service][sdk-grpc] accessible from outside the
|
||||
application, allowing transactions to be broadcast to the network, look up
|
||||
transactions, and simulate transaction costs.
|
||||
|
||||
|
||||
### RPC Transport
|
||||
|
||||
The consensus node RPC service allows callers to query consensus parameters
|
||||
(genesis data, transactions, commits), node status (network info, health
|
||||
checks), application state (abci_query, abci_info), mempool state, and other
|
||||
attributes of the node and its application. The service also provides methods
|
||||
allowing transactions and evidence to be injected ("broadcast") into the
|
||||
blockchain.
|
||||
|
||||
The RPC service is exposed in several ways:
|
||||
|
||||
- HTTP GET: Queries may be sent as URI parameters, with method names in the path.
|
||||
|
||||
- HTTP POST: Queries may be sent as JSON-RPC request messages in the body of an
|
||||
HTTP POST request. The server uses a custom implementation of JSON-RPC that
|
||||
is not fully compatible with the [JSON-RPC 2.0 spec][json-rpc], but handles
|
||||
the common cases.
|
||||
|
||||
- Websocket: Queries may be sent as JSON-RPC request messages via a websocket.
|
||||
This transport uses more or less the same JSON-RPC plumbing as the HTTP POST
|
||||
handler.
|
||||
|
||||
The websocket endpoint also includes three methods that are _only_ exported
|
||||
via websocket, which appear to support event subscription.
|
||||
|
||||
- gRPC: A subset of queries may be issued in protocol buffer format to the gRPC
|
||||
interface described above under (4). As noted, this endpoint is deprecated
|
||||
and will be removed in v0.36.
|
||||
|
||||
### Opportunities for Simplification
|
||||
|
||||
**Claim:** There are too many IPC mechanisms.
|
||||
|
||||
The preponderance of ABCI usage is via the Cosmos SDK, which means the
|
||||
application and the consensus node are compiled together into a single binary,
|
||||
and the consensus node calls the ABCI methods of the application directly as Go
|
||||
functions.
|
||||
|
||||
We also need a true IPC transport to support ABCI applications _not_ written in
|
||||
Go. There are also several known applications written in Rust, for example
|
||||
(including [Anoma](https://github.com/anoma/anoma), Penumbra,
|
||||
[Oasis](https://github.com/oasisprotocol/oasis-core), Twilight, and
|
||||
[Nomic](https://github.com/nomic-io/nomic)). Ideally we will have at most one
|
||||
such transport "built-in": More esoteric cases can be handled by a custom proxy.
|
||||
Pragmatically, gRPC is probably the right choice here.
|
||||
|
||||
The primary consumers of the multi-headed "RPC service" today are the light
|
||||
client and the `tendermint` command-line client. There is probably some local
|
||||
use via curl, but I expect that is mostly ad hoc. Ethan reports that nodes are
|
||||
often configured with the ports to the RPC service blocked, which is good for
|
||||
security but complicates use by the light client.
|
||||
|
||||
### Context: Remote Signer Issues
|
||||
|
||||
Since the remote signer needs a secure communication channel to exchange keys
|
||||
and signatures, and is expected to run truly remotely from the node (i.e., on a
|
||||
separate physical server), there is not a whole lot we can do here. We should
|
||||
finish the deprecation and removal of the "raw" socket protocol between the
|
||||
consensus node and remote signers, but the use of gRPC is appropriate.
|
||||
|
||||
The main improvement we can make is to simplify the implementation quite a bit,
|
||||
once we no longer need to support both "raw" and gRPC transports.
|
||||
|
||||
### Context: ABCI Issues
|
||||
|
||||
In the original design of ABCI, the presumption was that all access to the
|
||||
application should be mediated by the consensus node. The idea is that outside
|
||||
access could change application state and corrupt the consensus process, which
|
||||
relies on the application to be deterministic. Of course, even without outside
|
||||
access an application could behave nondeterministically, but allowing other
|
||||
programs to send it requests was seen as courting trouble.
|
||||
|
||||
Conversely, users noted that most of the time, tools written for a particular
|
||||
application don't want to talk to the consensus module directly. The
|
||||
application "owns" the state machine the consensus engine is replicating, so
|
||||
tools that care about application state should talk to the application.
|
||||
Otherwise, they would have to bake in knowledge about Tendermint (e.g., its
|
||||
interfaces and data structures) just because of the mediation.
|
||||
|
||||
For clients to talk directly to the application, however, there is another
|
||||
concern: The consensus node is the ABCI _client_, so it is inconvenient for the
|
||||
application to "push" work into the consensus module via ABCI itself. The
|
||||
current implementation works around this by calling the consensus node's RPC
|
||||
service, which exposes an `ABCIQuery` kitchen-sink method that allows the
|
||||
application a way to poke ABCI messages in the other direction.
|
||||
|
||||
Without this RPC method, you could work around this (at least in principle) by
|
||||
having the consensus module "poll" the application for work that needs done,
|
||||
but that has unsatisfactory implications for performance and robustness, as
|
||||
well as being harder to understand.
|
||||
|
||||
There has apparently been discussion about trying to make a more bidirectional
|
||||
communication between the consensus node and the application, but this issue
|
||||
seems to still be unresolved.
|
||||
|
||||
Another complication of ABCI is that it requires the application (server) to
|
||||
maintain [four separate connections][abci-conn]: One for "consensus" operations
|
||||
(BeginBlock, EndBlock, DeliverTx, Commit), one for "mempool" operations, one
|
||||
for "query" operations, and one for "snapshot" (state synchronization) operations.
|
||||
The rationale seems to have been that these groups of operations should be able
|
||||
to proceed concurrently with each other. In practice, it results in a very complex
|
||||
state management problem to coordinate state updates between the separate streams.
|
||||
While application authors in Go are mostly insulated from that complexity by the
|
||||
Cosmos SDK, the plumbing to maintain those separate streams is complicated, hard
|
||||
to understand, and we suspect it contains concurrency bugs and/or lock contention
|
||||
issues affecting performance that are subtle and difficult to pin down.
|
||||
|
||||
Even without changing the semantics of any ABCI operations, this code could be
|
||||
made smaller and easier to debug by separating the management of concurrency
|
||||
and locking from the IPC transport: If all requests and responses are routed
|
||||
through one connection, the server can explicitly maintain priority queues for
|
||||
requests and responses, and make less-conservative decisions about when locks
|
||||
are (or aren't) required to synchronize state access. With independent queues,
|
||||
the server must lock conservatively, and no optimistic scheduling is practical.
|
||||
|
||||
This would be a tedious implementation change, but should be achievable without
|
||||
breaking any of the existing interfaces. More importantly, it could potentially
|
||||
address a lot of difficult concurrency and performance problems we currently
|
||||
see anecdotally but have difficultly isolating because of how intertwined these
|
||||
separate message streams are at runtime.
|
||||
|
||||
TODO: Impact of ABCI++ for this topic?
|
||||
|
||||
### Context: RPC Issues
|
||||
|
||||
The RPC system serves several masters, and has a complex surface area. I
|
||||
believe there are some improvements that can be exposed by separating some of
|
||||
these concerns.
|
||||
|
||||
The Tendermint light client currently uses the RPC service to look up blocks
|
||||
and transactions, and to forward ABCI queries to the application. The light
|
||||
client proxy uses the RPC service via a websocket. The Cosmos IBC relayer also
|
||||
uses the RPC service via websocket to watch for transaction events, and uses
|
||||
the `ABCIQuery` method to fetch information and proofs for posted transactions.
|
||||
|
||||
Some work is already underway toward using P2P message passing rather than RPC
|
||||
to synchronize light client state with the rest of the network. IBC relaying,
|
||||
however, requires access to the event system, which is currently not accessible
|
||||
except via the RPC interface. Event subscription _could_ be exposed via P2P,
|
||||
but that is a larger project since it adds P2P communication load, and might
|
||||
thus have an impact on the performance of consensus.
|
||||
|
||||
If event subscription can be moved into the P2P network, we could entirely
|
||||
remove the websocket transport, even for clients that still need access to the
|
||||
RPC service. Until then, we may still be able to reduce the scope of the
|
||||
websocket endpoint to _only_ event subscription, by moving uses of the RPC
|
||||
server as a proxy to ABCI over to the gRPC interface.
|
||||
|
||||
Having the RPC server still makes sense for local bootstrapping and operations,
|
||||
but can be further simplified. Here are some specific proposals:
|
||||
|
||||
- Remove the HTTP GET interface entirely.
|
||||
|
||||
- Simplify JSON-RPC plumbing to remove unnecessary reflection and wrapping.
|
||||
|
||||
- Remove the gRPC interface (this is already planned for v0.36).
|
||||
|
||||
- Separate the websocket interface from the rest of the RPC service, and
|
||||
restrict it to only event subscription.
|
||||
|
||||
Eventually we should try to emove the websocket interface entirely, but we
|
||||
will need to revisit that (probably in a new RFC) once we've done some of the
|
||||
easier things.
|
||||
|
||||
These changes would preserve the ability of operators to issue queries with
|
||||
curl (but would require using JSON-RPC instead of URI parameters). That would
|
||||
be a little less user-friendly, but for a use case that should not be that
|
||||
prevalent.
|
||||
|
||||
These changes would also preserve compatibility with existing JSON-RPC based
|
||||
code paths like the `tendermint` CLI and the light client (even ahead of
|
||||
further work to remove that dependency).
|
||||
|
||||
**Design goal:** An operator should be able to disable non-local access to the
|
||||
RPC server on any node in the network without impairing the ability of the
|
||||
network to function for service of state replication, including light clients.
|
||||
|
||||
**Design principle:** All communication required to implement and monitor the
|
||||
consensus network should use P2P, including the various synchronizations.
|
||||
|
||||
### Options for ABCI Transport
|
||||
|
||||
The majority of current usage is in Go, and the majority of that is mediated by
|
||||
the Cosmos SDK, which uses the "direct call" interface. There is probably some
|
||||
opportunity to clean up the implementation of that code, notably by inverting
|
||||
which interface is at the "top" of the abstraction stack (currently it acts
|
||||
like an RPC interface, and escape-hatches into the direct call). However, this
|
||||
general approach works fine and doesn't need to be fundamentally changed.
|
||||
|
||||
For applications _not_ written in Go, the two remaining options are the
|
||||
"socket" protocol (another variation on varint-prefixed protobuf messages over
|
||||
an unstructured stream) and gRPC. It would be nice if we could get rid of one
|
||||
of these to reduce (unneeded?) optionality.
|
||||
|
||||
Since both the socket protocol and gRPC depend on protocol buffers, the
|
||||
"socket" protocol is the most obvious choice to remove. While gRPC is more
|
||||
complex, the set of languages that _have_ protobuf support but _lack_ gRPC
|
||||
support is small. Moreover, gRPC is already widely used in the rest of the
|
||||
ecosystem (including the Cosmos SDK).
|
||||
|
||||
If some use case did arise later that can't work with gRPC, it would not be too
|
||||
difficult for that application author to write a little proxy (in Go) that
|
||||
bridges the convenient SDK APIs into a simpler protocol than gRPC.
|
||||
|
||||
**Design principle:** It is better for an uncommon special case to carry the
|
||||
burdens of its specialness, than to bake an escape hatch into the infrastructure.
|
||||
|
||||
**Recommendation:** We should deprecate and remove the socket protocol.
|
||||
|
||||
### Options for RPC Transport
|
||||
|
||||
[ADR 057][adr-57] proposes using gRPC for the Tendermint RPC implementation.
|
||||
This is still possible, but if we are able to simplify and decouple the
|
||||
concerns as described above, I do not think it should be necessary.
|
||||
|
||||
While JSON-RPC is not the best possible RPC protocol for all situations, it has
|
||||
some advantages over gRPC for our domain. Specifically:
|
||||
|
||||
- It is easy to call JSON-RPC manually from the command-line, which helps with
|
||||
a common concern for the RPC service, local debugging and operations.
|
||||
|
||||
Relatedly: JSON is relatively easy for humans to read and write, and it can
|
||||
be easily copied and pasted to share sample queries and debugging results in
|
||||
chat, issue comments, and so on. Ideally, the RPC service will not be used
|
||||
for activities where the costs of a text protocol are important compared to
|
||||
its legibility and manual usability benefits.
|
||||
|
||||
- gRPC has an enormous dependency footprint for both clients and servers, and
|
||||
many of the features it provides to support security and performance
|
||||
(encryption, compression, streaming, etc.) are mostly irrelevant to local
|
||||
use. Tendermint already needs to include a gRPC client for the remote signer,
|
||||
but if we can avoid the need for a _client_ to depend on gRPC, that is a win
|
||||
for usability.
|
||||
|
||||
- If we intend to migrate light clients off RPC to use P2P entirely, there is
|
||||
no advantage to forcing a temporary migration to gRPC along the way; and once
|
||||
the light client is not dependent on the RPC service, the efficiency of the
|
||||
protocol is much less important.
|
||||
|
||||
- We can still get the benefits of generated data types using protocol buffers, even
|
||||
without using gRPC:
|
||||
|
||||
- Protobuf defines a standard JSON encoding for all message types so
|
||||
languages with protobuf support do not need to worry about type mapping
|
||||
oddities.
|
||||
|
||||
- Using JSON means that even languages _without_ good protobuf support can
|
||||
implement the protocol with a bit more work, and I expect this situation to
|
||||
be rare.
|
||||
|
||||
Even if a language lacks a good standard JSON-RPC mechanism, the protocol is
|
||||
lightweight and can be implemented by simple send/receive over TCP or
|
||||
Unix-domain sockets with no need for code generation, encryption, etc. gRPC
|
||||
uses a complex HTTP/2 based transport that is not easily replicated.
|
||||
|
||||
### Future Work
|
||||
|
||||
The background and proposals sketched above focus on the existing structure of
|
||||
Tendermint and improvements we can make in the short term. It is worthwhile to
|
||||
also consider options for longer-term broader changes to the IPC ecosystem.
|
||||
The following outlines some ideas at a high level:
|
||||
|
||||
- **Consensus service:** Today, the application and the consensus node are
|
||||
nominally connected only via ABCI. Tendermint was originally designed with
|
||||
the assumption that all communication with the application should be mediated
|
||||
by the consensus node. Based on further experience, however, the design goal
|
||||
is now that the _application_ should be the mediator of application state.
|
||||
|
||||
As noted above, however, ABCI is a client/server protocol, with the
|
||||
application as the server. For outside clients that turns out to have been a
|
||||
good choice, but it complicates the relationship between the application and
|
||||
the consensus node: Previously transactions were entered via the node, now
|
||||
they are entered via the app.
|
||||
|
||||
We have worked around this by using the Tendermint RPC service to give the
|
||||
application a "back channel" to the consensus node, so that it can push
|
||||
transactions back into the consensus network. But the RPC service exposes a
|
||||
lot of other functionality, too, including event subscription, block and
|
||||
transaction queries, and a lot of node status information.
|
||||
|
||||
Even if we can't easily "fix" the orientation of the ABCI relationship, we
|
||||
could improve isolation by splitting out the parts of the RPC service that
|
||||
the application needs as a back-channel, and sharing those _only_ with the
|
||||
application. By defining a "consensus service", we could give the application
|
||||
a way to talk back limited to only the capabilities it needs. This approach
|
||||
has the benefit that we could do it without breaking existing use, and if we
|
||||
later did "fix" the ABCI directionality, we could drop the special case
|
||||
without disrupting the rest of the RPC interface.
|
||||
|
||||
- **Event service:** Right now, the IBC relayer relies on the Tendermint RPC
|
||||
service to provide a stream of block and transaction events, which it uses to
|
||||
discover which transactions need relaying to other chains. While I think
|
||||
that event subscription should eventually be handled via P2P, we could gain
|
||||
some immediate benefit by splitting out event subscription from the rest of
|
||||
the RPC service.
|
||||
|
||||
In this model, an event subscription service would be exposed on the public
|
||||
network, but on a different endpoint. This would remove the need for the RPC
|
||||
service to support the websocket protocol, and would allow operators to
|
||||
isolate potentially sensitive status query results from the public network.
|
||||
|
||||
At the moment the relayers also use the RPC service to get block data for
|
||||
synchronization, but work is already in progress to handle that concern via
|
||||
the P2P layer. Once that's done, event subscription could be separated.
|
||||
|
||||
Separating parts of the existing RPC service is not without cost: It might
|
||||
require additional connection endpoints, for example, though it is also not too
|
||||
difficult for multiple otherwise-independent services to share a connection.
|
||||
|
||||
In return, though, it would become easier to reduce transport options and for
|
||||
operators to independently control access to sensitive data. Considering the
|
||||
viability and implications of these ideas is beyond the scope of this RFC, but
|
||||
they are documented here since they follow from the background we have already
|
||||
discussed.
|
||||
|
||||
## References
|
||||
|
||||
[abci]: https://github.com/tendermint/tendermint/tree/master/spec/abci
|
||||
[rpc-service]: https://docs.tendermint.com/master/rpc/
|
||||
[light-client]: https://docs.tendermint.com/master/tendermint-core/light-client.html
|
||||
[tm-cli]: https://github.com/tendermint/tendermint/tree/master/cmd/tendermint
|
||||
[cosmos-sdk]: https://github.com/cosmos/cosmos-sdk/
|
||||
[local-client]: https://github.com/tendermint/tendermint/blob/master/abci/client/local_client.go
|
||||
[socket-server]: https://github.com/tendermint/tendermint/blob/master/abci/server/socket_server.go
|
||||
[sdk-grpc]: https://pkg.go.dev/github.com/cosmos/cosmos-sdk/types/tx#ServiceServer
|
||||
[json-rpc]: https://www.jsonrpc.org/specification
|
||||
[abci-conn]: https://github.com/tendermint/tendermint/blob/master/spec/abci/apps.md#state
|
||||
[adr-57]: https://github.com/tendermint/tendermint/blob/master/docs/architecture/adr-057-RPC.md
|
||||
@@ -0,0 +1,283 @@
|
||||
# RFC 003: Taxonomy of potential performance issues in Tendermint
|
||||
|
||||
## Changelog
|
||||
|
||||
- 2021-09-02: Created initial draft (@wbanfield)
|
||||
- 2021-09-14: Add discussion of the event system (@wbanfield)
|
||||
|
||||
## Abstract
|
||||
|
||||
This document discusses the various sources of performance issues in Tendermint and
|
||||
attempts to clarify what work may be required to understand and address them.
|
||||
|
||||
## Background
|
||||
|
||||
Performance, loosely defined as the ability of a software process to perform its work
|
||||
quickly and efficiently under load and within reasonable resource limits, is a frequent
|
||||
topic of discussion in the Tendermint project.
|
||||
To effectively address any issues with Tendermint performance we need to
|
||||
categorize the various issues, understand their potential sources, and gauge their
|
||||
impact on users.
|
||||
|
||||
Categorizing the different known performance issues will allow us to discuss and fix them
|
||||
more systematically. This document proposes a rough taxonomy of performance issues
|
||||
and highlights areas where more research into potential performance problems is required.
|
||||
|
||||
Understanding Tendermint's performance limitations will also be critically important
|
||||
as we make changes to many of its subsystems. Performance is a central concern for
|
||||
upcoming decisions regarding the `p2p` protocol, RPC message encoding and structure,
|
||||
database usage and selection, and consensus protocol updates.
|
||||
|
||||
|
||||
## Discussion
|
||||
|
||||
This section attempts to delineate the different sections of Tendermint functionality
|
||||
that are often cited as having performance issues. It raises questions and suggests
|
||||
lines of inquiry that may be valuable for better understanding Tendermint's performance issues.
|
||||
|
||||
As a note: We should avoid quickly adding many microbenchmarks or package level benchmarks.
|
||||
These are prone to being worse than useless as they can obscure what _should_ be
|
||||
focused on: performance of the system from the perspective of a user. We should,
|
||||
instead, tune performance with an eye towards user needs and actions users make. These users comprise
|
||||
both operators of Tendermint chains and the people generating transactions for
|
||||
Tendermint chains. Both of these sets of users are largely aligned in wanting an end-to-end
|
||||
system that operates quickly and efficiently.
|
||||
|
||||
REQUEST: The list below may be incomplete, if there are additional sections that are often
|
||||
cited as creating poor performance, please comment so that they may be included.
|
||||
|
||||
### P2P
|
||||
|
||||
#### Claim: Tendermint cannot scale to large numbers of nodes
|
||||
|
||||
A complaint has been reported that Tendermint networks cannot scale to large numbers of nodes.
|
||||
The listed number of nodes a user reported as causing issue was in the thousands.
|
||||
We don't currently have evidence about what the upper-limit of nodes that Tendermint's
|
||||
P2P stack can scale to.
|
||||
|
||||
We need to more concretely understand the source of issues and determine what layer
|
||||
is causing a problem. It's possible that the P2P layer, in the absence of any reactors
|
||||
sending data, is perfectly capable of managing thousands of peer connections. For
|
||||
a reasonable networking and application setup, thousands of connections should not present any
|
||||
issue for the application.
|
||||
|
||||
We need more data to understand the problem directly. We want to drive the popularity
|
||||
and adoption of Tendermint and this will mean allowing for chains with more validators.
|
||||
We should follow up with users experiencing this issue. We may then want to add
|
||||
a series of metrics to the P2P layer to better understand the inefficiencies it produces.
|
||||
|
||||
The following metrics can help us understand the sources of latency in the Tendermint P2P stack:
|
||||
|
||||
* Number of messages sent and received per second
|
||||
* Time of a message spent on the P2P layer send and receive queues
|
||||
|
||||
The following metrics exist and should be leveraged in addition to those added:
|
||||
|
||||
* Number of peers node's connected to
|
||||
* Number of bytes per channel sent and received from each peer
|
||||
|
||||
### Sync
|
||||
|
||||
#### Claim: Block Syncing is slow
|
||||
|
||||
Bootstrapping a new node in a network to the height of the rest of the network is believed to
|
||||
take longer than users would like. Block sync requires fetching all of the blocks from
|
||||
peers and placing them into the local disk for storage. A useful line of inquiry
|
||||
is understanding how quickly a perfectly tuned system _could_ fetch all of the state
|
||||
over a network so that we understand how much overhead Tendermint actually adds.
|
||||
|
||||
The operation is likely to be _incredibly_ dependent on the environment in which
|
||||
the node is being run. The factors that will influence syncing include:
|
||||
1. Number of peers that a syncing node may fetch from.
|
||||
2. Speed of the disk that a validator is writing to.
|
||||
3. Speed of the network connection between the different peers that node is
|
||||
syncing from.
|
||||
|
||||
We should calculate how quickly this operation _could possibly_ complete for common chains and nodes.
|
||||
To calculate how quickly this operation could possibly complete, we should assume that
|
||||
a node is reading at line-rate of the NIC and writing at the full drive speed to its
|
||||
local storage. Comparing this theoretical upper-limit to the actual sync times
|
||||
observed by node operators will give us a good point of comparison for understanding
|
||||
how much overhead Tendermint incurs.
|
||||
|
||||
We should additionally add metrics to the blocksync operation to more clearly pinpoint
|
||||
slow operations. The following metrics should be added to the block syncing operation:
|
||||
|
||||
* Time to fetch and validate each block
|
||||
* Time to execute a block
|
||||
* Blocks sync'd per unit time
|
||||
|
||||
### Application
|
||||
|
||||
Applications performing complex state transitions have the potential to bottleneck
|
||||
the Tendermint node.
|
||||
|
||||
#### Claim: ABCI block delivery could cause slowdown
|
||||
|
||||
ABCI delivers blocks in several methods: `BeginBlock`, `DeliverTx`, `EndBlock`, `Commit`.
|
||||
|
||||
Tendermint delivers transactions one-by-one via the `DeliverTx` call. Most of the
|
||||
transaction delivery in Tendermint occurs asynchronously and therefore appears unlikely to
|
||||
form a bottleneck in ABCI.
|
||||
|
||||
After delivering all transactions, Tendermint then calls the `Commit` ABCI method.
|
||||
Tendermint [locks all access to the mempool][abci-commit-description] while `Commit`
|
||||
proceeds. This means that an application that is slow to execute all of its
|
||||
transactions or finalize state during the `Commit` method will prevent any new
|
||||
transactions from being added to the mempool. Apps that are slow to commit will
|
||||
prevent consensus from proceeded to the next consensus height since Tendermint
|
||||
cannot validate block proposals or produce block proposals without the
|
||||
AppHash obtained from the `Commit` method. We should add a metric for each
|
||||
step in the ABCI protocol to track the amount of time that a node spends communicating
|
||||
with the application at each step.
|
||||
|
||||
#### Claim: ABCI serialization overhead causes slowdown
|
||||
|
||||
The most common way to run a Tendermint application is using the Cosmos-SDK.
|
||||
The Cosmos-SDK runs the ABCI application within the same process as Tendermint.
|
||||
When an application is run in the same process as Tendermint, a serialization penalty
|
||||
is not paid. This is because the local ABCI client does not serialize method calls
|
||||
and instead passes the protobuf type through directly. This can be seen
|
||||
in [local_client.go][abci-local-client-code].
|
||||
|
||||
Serialization and deserialization in the gRPC and socket protocol ABCI methods
|
||||
may cause slowdown. While these may cause issue, they are not part of the primary
|
||||
usecase of Tendermint and do not necessarily need to be addressed at this time.
|
||||
|
||||
### RPC
|
||||
|
||||
#### Claim: The Query API is slow.
|
||||
|
||||
The query API locks a mutex across the ABCI connections. This causes consensus to
|
||||
slow during queries, as ABCI is no longer able to make progress. This is known
|
||||
to be causing issue in the cosmos-sdk and is being addressed [in the sdk][sdk-query-fix]
|
||||
but a more robust solution may be required. Adding metrics to each ABCI client connection
|
||||
and message as described in the Application section of this document would allow us
|
||||
to further introspect the issue here.
|
||||
|
||||
#### Claim: RPC Serialization may cause slowdown
|
||||
|
||||
The Tendermint RPC uses a modified version of JSON-RPC. This RPC powers the `broadcast_tx_*` methods,
|
||||
which is a critical method for adding transactions to Tendermint at the moment. This method is
|
||||
likely invoked quite frequently on popular networks. Being able to perform efficiently
|
||||
on this common and critical operation is very important. The current JSON-RPC implementation
|
||||
relies heavily on type introspection via reflection, which is known to be very slow in
|
||||
Go. We should therefore produce benchmarks of this method to determine how much overhead
|
||||
we are adding to what, is likely to be, a very common operation.
|
||||
|
||||
The other JSON-RPC methods are much less critical to the core functionality of Tendermint.
|
||||
While there may other points of performance consideration within the RPC, methods that do not
|
||||
receive high volumes of requests should not be prioritized for performance consideration.
|
||||
|
||||
NOTE: Previous discussion of the RPC framework was done in [ADR 57][adr-57] and
|
||||
there is ongoing work to inspect and alter the JSON-RPC framework in [RFC 002][rfc-002].
|
||||
Much of these RPC-related performance considerations can either wait until the work of RFC 002 work is done or be
|
||||
considered concordantly with the in-flight changes to the JSON-RPC.
|
||||
|
||||
### Protocol
|
||||
|
||||
#### Claim: Gossiping messages is a slow process
|
||||
|
||||
Currently, for any validator to successfully vote in a consensus _step_, it must
|
||||
receive votes from greater than 2/3 of the validators on the network. In many cases,
|
||||
it's preferable to receive as many votes as possible from correct validators.
|
||||
|
||||
This produces a quadratic increase in messages that are communicated as more validators join the network.
|
||||
(Each of the N validators must communicate with all other N-1 validators).
|
||||
|
||||
This large number of messages communicated per step has been identified to impact
|
||||
performance of the protocol. Given that the number of messages communicated has been
|
||||
identified as a bottleneck, it would be extremely valuable to gather data on how long
|
||||
it takes for popular chains with many validators to gather all votes within a step.
|
||||
|
||||
Metrics that would improve visibility into this include:
|
||||
|
||||
* Amount of time for a node to gather votes in a step.
|
||||
* Amount of time for a node to gather all block parts.
|
||||
* Number of votes each node sends to gossip (i.e. not its own votes, but votes it is
|
||||
transmitting for a peer).
|
||||
* Total number of votes each node sends to receives (A node may receive duplicate votes
|
||||
so understanding how frequently this occurs will be valuable in evaluating the performance
|
||||
of the gossip system).
|
||||
|
||||
#### Claim: Hashing Txs causes slowdown in Tendermint
|
||||
|
||||
Using a faster hash algorithm for Tx hashes is currently a point of discussion
|
||||
in Tendermint. Namely, it is being considered as part of the [modular hashing proposal][modular-hashing].
|
||||
It is currently unknown if hashing transactions in the Mempool forms a significant bottleneck.
|
||||
Although it does not appear to be documented as slow, there are a few open github
|
||||
issues that indicate a possible user preference for a faster hashing algorithm,
|
||||
including [issue 2187][issue-2187] and [issue 2186][issue-2186].
|
||||
|
||||
It is likely worth investigating what order of magnitude Tx hashing takes in comparison to other
|
||||
aspects of adding a Tx to the mempool. It is not currently clear if the rate of adding Tx
|
||||
to the mempool is a source of user pain. We should not endeavor to make large changes to
|
||||
consensus critical components without first being certain that the change is highly
|
||||
valuable and impactful.
|
||||
|
||||
### Digital Signatures
|
||||
|
||||
#### Claim: Verification of digital signatures may cause slowdown in Tendermint
|
||||
|
||||
Working with cryptographic signatures can be computationally expensive. The cosmos
|
||||
hub uses [ed25519 signatures][hub-signature]. The library performing signature
|
||||
verification in Tendermint on votes is [benchmarked][ed25519-bench] to be able to perform an `ed25519`
|
||||
signature in 75μs on a decently fast CPU. A validator in the Cosmos Hub performs
|
||||
3 sets of verifications on the signatures of the 140 validators in the Hub
|
||||
in a consensus round, during block verification, when verifying the prevotes, and
|
||||
when verifying the precommits. With no batching, this would be roughly `3ms` per
|
||||
round. It is quite unlikely, therefore, that this accounts for any serious amount
|
||||
of the ~7 seconds of block time per height in the Hub.
|
||||
|
||||
This may cause slowdown when syncing, since the process needs to constantly verify
|
||||
signatures. It's possible that improved signature aggregation will lead to improved
|
||||
light client or other syncing performance. In general, a metric should be added
|
||||
to track block rate while blocksyncing.
|
||||
|
||||
#### Claim: Our use of digital signatures in the consensus protocol contributes to performance issue
|
||||
|
||||
Currently, Tendermint's digital signature verification requires that all validators
|
||||
receive all vote messages. Each validator must receive the complete digital signature
|
||||
along with the vote message that it corresponds to. This means that all N validators
|
||||
must receive messages from at least 2/3 of the N validators in each consensus
|
||||
round. Given the potential for oddly shaped network topologies and the expected
|
||||
variable network roundtrip times of a few hundred milliseconds in a blockchain,
|
||||
it is highly likely that this amount of gossiping is leading to a significant amount
|
||||
of the slowdown in the Cosmos Hub and in Tendermint consensus.
|
||||
|
||||
### Tendermint Event System
|
||||
|
||||
#### Claim: The event system is a bottleneck in Tendermint
|
||||
|
||||
The Tendermint Event system is used to communicate and store information about
|
||||
internal Tendermint execution. The system uses channels internally to send messages
|
||||
to different subscribers. Sending an event [blocks on the internal channel][event-send].
|
||||
The default configuration is to [use an unbuffered channel for event publishes][event-buffer-capacity].
|
||||
Several consumers of the event system also use an unbuffered channel for reads.
|
||||
An example of this is the [event indexer][event-indexer-unbuffered], which takes an
|
||||
unbuffered subscription to the event system. The result is that these unbuffered readers
|
||||
can cause writes to the event system to block or slow down depending on contention in the
|
||||
event system. This has implications for the consensus system, which [publishes events][consensus-event-send].
|
||||
To better understand the performance of the event system, we should add metrics to track the timing of
|
||||
event sends. The following metrics would be a good start for tracking this performance:
|
||||
|
||||
* Time in event send, labeled by Event Type
|
||||
* Time in event receive, labeled by subscriber
|
||||
* Event throughput, measured in events per unit time.
|
||||
|
||||
### References
|
||||
[modular-hashing]: https://github.com/tendermint/tendermint/pull/6773
|
||||
[issue-2186]: https://github.com/tendermint/tendermint/issues/2186
|
||||
[issue-2187]: https://github.com/tendermint/tendermint/issues/2187
|
||||
[rfc-002]: https://github.com/tendermint/tendermint/pull/6913
|
||||
[adr-57]: https://github.com/tendermint/tendermint/blob/master/docs/architecture/adr-057-RPC.md
|
||||
[issue-1319]: https://github.com/tendermint/tendermint/issues/1319
|
||||
[abci-commit-description]: https://github.com/tendermint/tendermint/blob/master/spec/abci/apps.md#commit
|
||||
[abci-local-client-code]: https://github.com/tendermint/tendermint/blob/511bd3eb7f037855a793a27ff4c53c12f085b570/abci/client/local_client.go#L84
|
||||
[hub-signature]: https://github.com/cosmos/gaia/blob/0ecb6ed8a244d835807f1ced49217d54a9ca2070/docs/resources/genesis.md#consensus-parameters
|
||||
[ed25519-bench]: https://github.com/oasisprotocol/curve25519-voi/blob/d2e7fc59fe38c18ca990c84c4186cba2cc45b1f9/PERFORMANCE.md
|
||||
[event-send]: https://github.com/tendermint/tendermint/blob/5bd3b286a2b715737f6d6c33051b69061d38f8ef/libs/pubsub/pubsub.go#L338
|
||||
[event-buffer-capacity]: https://github.com/tendermint/tendermint/blob/5bd3b286a2b715737f6d6c33051b69061d38f8ef/types/event_bus.go#L14
|
||||
[event-indexer-unbuffered]: https://github.com/tendermint/tendermint/blob/5bd3b286a2b715737f6d6c33051b69061d38f8ef/state/indexer/indexer_service.go#L39
|
||||
[consensus-event-send]: https://github.com/tendermint/tendermint/blob/5bd3b286a2b715737f6d6c33051b69061d38f8ef/internal/consensus/state.go#L1573
|
||||
[sdk-query-fix]: https://github.com/cosmos/cosmos-sdk/pull/10045
|
||||
@@ -0,0 +1,213 @@
|
||||
========================================
|
||||
RFC 004: E2E Test Framework Enhancements
|
||||
========================================
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2021-09-14: started initial draft (@tychoish)
|
||||
|
||||
Abstract
|
||||
--------
|
||||
|
||||
This document discusses a series of improvements to the e2e test framework
|
||||
that we can consider during the next few releases to help boost confidence in
|
||||
Tendermint releases, and improve developer efficiency.
|
||||
|
||||
Background
|
||||
----------
|
||||
|
||||
During the 0.35 release cycle, the E2E tests were a source of great
|
||||
value, helping to identify a number of bugs before release. At the same time,
|
||||
the tests were not consistently passing during this time, thereby reducing
|
||||
their value, and forcing the core development team to allocate time and energy
|
||||
to maintaining and chasing down issues with the e2e tests and the test
|
||||
harness. The experience of this release cycle calls to mind a series of
|
||||
improvements to the test framework, and this document attempts to capture
|
||||
these improvements, along with motivations, and potential for impact.
|
||||
|
||||
Projects
|
||||
--------
|
||||
|
||||
Flexible Workload Generation
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Presently the e2e suite contains a single workload generation pattern, which
|
||||
exists simply to ensure that the test networks have some work during their
|
||||
runs. However, the shape and volume of the work is very consistent and is very
|
||||
gentle to help ensure test reliability.
|
||||
|
||||
We don't need a complex workload generation framework, but being able to have
|
||||
a few different workload shapes available for test networks, both generated and
|
||||
hand-crafted, would be useful.
|
||||
|
||||
Workload patterns/configurations might include:
|
||||
|
||||
- transaction targeting patterns (include light nodes, round robin, target
|
||||
individual nodes)
|
||||
|
||||
- variable transaction size over time.
|
||||
|
||||
- transaction broadcast option (synchronously, checked, fire-and-forget,
|
||||
mixed).
|
||||
|
||||
- number of transactions to submit.
|
||||
|
||||
- non-transaction workloads: (evidence submission, query, event subscription.)
|
||||
|
||||
Configurable Generator
|
||||
~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
The nightly e2e suite is defined by the `testnet generator
|
||||
<https://github.com/tendermint/tendermint/blob/master/test/e2e/generator/generate.go#L13-L65>`_,
|
||||
and it's difficult to add dimensions or change the focus of the test suite in
|
||||
any way without modifying the implementation of the generator. If the
|
||||
generator were more configurable, potentially via a file rather than in
|
||||
the Go implementation, we could modify the focus of the test suite on the
|
||||
fly.
|
||||
|
||||
Features that we might want to configure:
|
||||
|
||||
- number of test networks to generate of various topologies, to improve
|
||||
coverage of different configurations.
|
||||
|
||||
- test application configurations (to modify the latency of ABCI calls, etc.)
|
||||
|
||||
- size of test networks.
|
||||
|
||||
- workload shape and behavior.
|
||||
|
||||
- initial sync and catch-up configurations.
|
||||
|
||||
The workload generator currently provides runtime options for limiting the
|
||||
generator to specific types of P2P stacks, and for generating multiple groups
|
||||
of test cases to support parallelism. The goal is to extend this pattern and
|
||||
avoid hardcoding the matrix of test cases in the generator code. Once the
|
||||
testnet configuration generation behavior is configurable at runtime,
|
||||
developers may be able to use the e2e framework to validate changes before
|
||||
landing changes that break e2e tests a day later.
|
||||
|
||||
In addition to the autogenerated suite, it might make sense to maintain a
|
||||
small collection of hand-crafted cases that exercise configurations of
|
||||
concern, to run as part of the nightly (or less frequent) loop.
|
||||
|
||||
Implementation Plan Structure
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
As a development team, we should determine the features should impact the e2e
|
||||
testing early in the development cycle, and if we intend to modify the e2e
|
||||
tests to exercise a feature, we should identify this early and begin the
|
||||
integration process as early as possible.
|
||||
|
||||
To facilitate this, we should adopt a practice whereby we exercise specific
|
||||
features that are currently under development more rigorously in the e2e
|
||||
suite, and then as development stabilizes we can reduce the number or weight
|
||||
of these features in the suite.
|
||||
|
||||
As of 0.35 there are essentially two end to end tests: the suite of 64
|
||||
generated test networks, and the hand crafted `ci.toml` test case. The
|
||||
generated test cases help provide systemtic coverage, while the `ci` run
|
||||
provides coverage for a large number of features.
|
||||
|
||||
Reduce Cycle Time
|
||||
~~~~~~~~~~~~~~~~~
|
||||
|
||||
One of the barriers to leveraging the e2e framework, and one of the challenges
|
||||
in debugging failures, is the cycle time of running a single test iteration is
|
||||
quite high: 5 minutes to build the docker image, plus the time to run the test
|
||||
or tests.
|
||||
|
||||
There are a number of improvements and enhancements that can reduce the cycle
|
||||
time in practice:
|
||||
|
||||
- reduce the amount of time required to build the docker image used in these
|
||||
tests. Without the dependency on CGo, the tendermint binaries could be
|
||||
(cross) compiled outside of the docker container and then injected into
|
||||
them, which would take better advantage of docker's native caching,
|
||||
although, without the dependency on CGo there would be no hard requirement
|
||||
for the e2e tests to use docker.
|
||||
|
||||
- support test parallelism. Because of the way the testnets are orchestrated
|
||||
a single system can really only run one network at a time. For executions
|
||||
(local or remote) with more resources, there's no reason to run a few
|
||||
networks in parallel to reduce the feedback time.
|
||||
|
||||
- prune testnet configurations that are unlikely to provide good signal, to
|
||||
shorten the time to feedback.
|
||||
|
||||
- apply some kind of tiered approach to test execution, to improve the
|
||||
legibility of the test result. For example order tests by the dependency of
|
||||
their features, or run test networks without perturbations before running
|
||||
that configuration with perturbations, to be able to isolate the impact of
|
||||
specific features.
|
||||
|
||||
- orchestrate the test harness directly from go test rather than via a special
|
||||
harness and shell scripts so e2e tests may more naively fit into developers
|
||||
existing workflows.
|
||||
|
||||
Many of these improvements, particularly, reducing the build time will also
|
||||
reduce the time to get feedback during automated builds.
|
||||
|
||||
Deeper Insights
|
||||
~~~~~~~~~~~~~~~
|
||||
|
||||
When a test network fails, it's incredibly difficult to understand _why_ the
|
||||
network failed, as the current system provides very little insight into the
|
||||
system outside of the process logs. When a test network stalls or fails
|
||||
developers should be able to quickly and easily get a sense of the state of
|
||||
the network and all nodes.
|
||||
|
||||
Improvements in persuit of this goal, include functionality that would help
|
||||
node operators in production environments by improving the quality and utility
|
||||
of the logging messages and other reported metrics, but also provide some
|
||||
tools to collect and aggregate this data for developers in the context of test
|
||||
networks.
|
||||
|
||||
- Interleave messages from all nodes in the network to be able to correlate
|
||||
events during the test run.
|
||||
|
||||
- Collect structured metrics of the system operation (CPU/MEM/IO) during the
|
||||
test run, as well as from each tendermint/application process.
|
||||
|
||||
- Build (simple) tools to be able to render and summarize the data collected
|
||||
during the test run to answer basic questions about test outcome.
|
||||
|
||||
Flexible Assertions
|
||||
~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Currently, all assertions run for every test network, which makes the
|
||||
assertions pretty bland, and the framework primarily useful as a smoke-test
|
||||
framework, but it might be useful to be able to write and run different
|
||||
tests for different configurations. This could allow us to test outside of the
|
||||
happy-path.
|
||||
|
||||
In general our existing assertions occupy a fraction of the total test time,
|
||||
so the relative cost of adding a few extra test assertions would be of limited
|
||||
cost, and could help build confidence.
|
||||
|
||||
Additional Kinds of Testing
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
The existing e2e suite, exercises networks of nodes that have homogeneous
|
||||
tendermint version, stable configuration, that are expected to make
|
||||
progress. There are many other possible test configurations that may be
|
||||
interesting to engage with. These could include dimensions, such as:
|
||||
|
||||
- Multi-version testing to exercise our compatibility guarantees for networks
|
||||
that might have different tendermint versions.
|
||||
|
||||
- As a flavor or mult-version testing, include upgrade testing, to build
|
||||
confidence in migration code and procedures.
|
||||
|
||||
- Additional test applications, particularly practical-type applciations
|
||||
including some that use gaiad and/or the cosmos-sdk. Test-only applications
|
||||
that simulate other kinds of applications (e.g. variable application
|
||||
operation latency.)
|
||||
|
||||
- Tests of "non-viable" configurations that ensure that forbidden combinations
|
||||
lead to halts.
|
||||
|
||||
References
|
||||
----------
|
||||
|
||||
- `ADR 66: End-to-End Testing <../architecture/adr-66-e2e-testing.md>`_
|
||||
@@ -0,0 +1,122 @@
|
||||
=====================
|
||||
RFC 005: Event System
|
||||
=====================
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2021-09-17: Initial Draft (@tychoish)
|
||||
|
||||
Abstract
|
||||
--------
|
||||
|
||||
The event system within Tendermint, which supports a lot of core
|
||||
functionality, also represents a major infrastructural liability. As part of
|
||||
our upcoming review of the RPC interfaces and our ongoing thoughts about
|
||||
stability and performance, as well as the preparation for Tendermint 1.0, we
|
||||
should revisit the design and implementation of the event system. This
|
||||
document discusses both the current state of the system and potential
|
||||
directions for future improvement.
|
||||
|
||||
Background
|
||||
----------
|
||||
|
||||
Current State of Events
|
||||
~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
The event system makes it possible for clients, both internal and external,
|
||||
to receive notifications of state replication events, such as new blocks,
|
||||
new transactions, validator set changes, as well as intermediate events during
|
||||
consensus. Because the event system is very cross cutting, the behavior and
|
||||
performance of the event publication and subscription system has huge impacts
|
||||
for all of Tendermint.
|
||||
|
||||
The subscription service is exposed over the RPC interface, but also powers
|
||||
the indexing (e.g. to an external database,) and is the mechanism by which
|
||||
`BroadcastTxCommit` is able to wait for transactions to land in a block.
|
||||
|
||||
The current pubsub mechanism relies on a couple of buffered channels,
|
||||
primarily between all event creators and subscribers, but also for each
|
||||
subscription. The result of this design is that, in some situations with the
|
||||
right collection of slow subscription consumers the event system can put
|
||||
backpressure on the consensus state machine and message gossiping in the
|
||||
network, thereby causing nodes to lag.
|
||||
|
||||
Improvements
|
||||
~~~~~~~~~~~~
|
||||
|
||||
The current system relies on implicit, bounded queues built by the buffered channels,
|
||||
and though threadsafe, can force all activity within Tendermint to serialize,
|
||||
which does not need to happen. Additionally, timeouts for subscription
|
||||
consumers related to the implementation of the RPC layer, may complicate the
|
||||
use of the system.
|
||||
|
||||
References
|
||||
~~~~~~~~~~
|
||||
|
||||
- Legacy Implementation
|
||||
- `publication of events <https://github.com/tendermint/tendermint/blob/master/libs/pubsub/pubsub.go#L333-L345>`_
|
||||
- `send operation <https://github.com/tendermint/tendermint/blob/master/libs/pubsub/pubsub.go#L489-L527>`_
|
||||
- `send loop <https://github.com/tendermint/tendermint/blob/master/libs/pubsub/pubsub.go#L381-L402>`_
|
||||
- Related RFCs
|
||||
- `RFC 002: IPC Ecosystem <./rfc-002-ipc-ecosystem.md>`_
|
||||
- `RFC 003: Performance Questions <./rfc-003-performance-questions.md>`_
|
||||
|
||||
Discussion
|
||||
----------
|
||||
|
||||
Changes to Published Events
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
As part of this process, the Tendermint team should do a study of the existing
|
||||
event types and ensure that there are viable production use cases for
|
||||
subscriptions to all event types. Instinctively it seems plausible that some
|
||||
of the events may not be useable outside of tendermint, (e.g. ``TimeoutWait``
|
||||
or ``NewRoundStep``) and it might make sense to remove them. Certainly, it
|
||||
would be good to make sure that we don't maintain infrastructure for unused or
|
||||
un-useful message indefinitely.
|
||||
|
||||
Blocking Subscription
|
||||
~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
The blocking subscription mechanism makes it possible to have *send*
|
||||
operations into the subscription channel be un-buffered (the event processing
|
||||
channel is still buffered.) In the blocking case, events from one subscription
|
||||
can block processing that event for other non-blocking subscriptions. The main
|
||||
case, it seems for blocking subscriptions is ensuring that a transaction has
|
||||
been committed to a block for ``BroadcastTxCommit``. Removing blocking
|
||||
subscriptions entirely, and potentially finding another way to implement
|
||||
``BroadcastTxCommit``, could lead to important simplifications and
|
||||
improvements to throughput without requiring large changes.
|
||||
|
||||
Subscription Identification
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Before `#6386 <https://github.com/tendermint/tendermint/pull/6386>`_, all
|
||||
subscriptions were identified by the combination of a client ID and a query,
|
||||
and with that change, it became possible to identify all subscription given
|
||||
only an ID, but compatibility with the legacy identification means that there's a
|
||||
good deal of legacy code as well as client side efficiency that could be
|
||||
improved.
|
||||
|
||||
Pubsub Changes
|
||||
~~~~~~~~~~~~~~
|
||||
|
||||
The pubsub core should be implemented in a way that removes the possibility of
|
||||
backpressure from the event system to impact the core system *or* for one
|
||||
subscription to impact the behavior of another area of the
|
||||
system. Additionally, because the current system is implemented entirely in
|
||||
terms of a collection of buffered channels, the event system (and large
|
||||
numbers of subscriptions) can be a source of memory pressure.
|
||||
|
||||
These changes could include:
|
||||
|
||||
- explicit cancellation and timeouts promulgated from callers (e.g. RPC end
|
||||
points, etc,) this should be done using contexts.
|
||||
|
||||
- subscription system should be able to spill to disk to avoid putting memory
|
||||
pressure on the core behavior of the node (consensus, gossip).
|
||||
|
||||
- subscriptions implemented as cursors rather than channels, with either
|
||||
condition variables to simulate the existing "push" API or a client side
|
||||
iterator API with some kind of long polling-type interface.
|
||||
@@ -0,0 +1,204 @@
|
||||
# RFC 006: Event Subscription
|
||||
|
||||
## Changelog
|
||||
|
||||
- 30-Oct-2021: Initial draft (@creachadair)
|
||||
|
||||
## Abstract
|
||||
|
||||
The Tendermint consensus node allows clients to subscribe to its event stream
|
||||
via methods on its RPC service. The ability to view the event stream is
|
||||
valuable for clients, but the current implementation has some deficiencies that
|
||||
make it difficult for some clients to use effectively. This RFC documents these
|
||||
issues and discusses possible approaches to solving them.
|
||||
|
||||
|
||||
## Background
|
||||
|
||||
A running Tendermint consensus node exports a [JSON-RPC service][rpc-service]
|
||||
that provides a [large set of methods][rpc-methods] for inspecting and
|
||||
interacting with the node. One important cluster of these methods are the
|
||||
`subscribe`, `unsubscribe`, and `unsubscribe_all` methods, which permit clients
|
||||
to subscribe to a filtered stream of the [events generated by the node][events]
|
||||
as it runs.
|
||||
|
||||
Unlike the other methods of the service, the methods in the "event
|
||||
subscription" cluster are not accessible via [ordinary HTTP GET or POST
|
||||
requests][rpc-transport], but require upgrading the HTTP connection to a
|
||||
[websocket][ws]. This is necessary because the `subscribe` request needs a
|
||||
persistent channel to deliver results back to the client, and an ordinary HTTP
|
||||
connection does not reliably persist across multiple requests. Since these
|
||||
methods do not work properly without a persistent channel, they are _only_
|
||||
exported via a websocket connection, and are not routed for plain HTTP.
|
||||
|
||||
|
||||
## Discussion
|
||||
|
||||
There are some operational problems with the current implementation of event
|
||||
subscription in the RPC service:
|
||||
|
||||
- **Event delivery is not valid JSON-RPC.** When a client issues a `subscribe`
|
||||
request, the server replies (correctly) with an initial empty acknowledgement
|
||||
(`{}`). After that, each matching event is delivered "unsolicited" (without
|
||||
another request from the client), as a separate [response object][json-response]
|
||||
with the same ID as the initial request.
|
||||
|
||||
This matters because it means a standard JSON-RPC client library can't
|
||||
interact correctly with the event subscription mechanism.
|
||||
|
||||
Even for clients that can handle unsolicited values pushed by the server,
|
||||
these responses are invalid: They have an ID, so they cannot be treated as
|
||||
[notifications][json-notify]; but the ID corresponds to a request that was
|
||||
already completed. In practice, this means that general-purpose JSON-RPC
|
||||
libraries cannot use this method correctly -- it requires a custom client.
|
||||
|
||||
The Go RPC client from the Tendermint core can support this case, but clients
|
||||
in other languages have no easy solution.
|
||||
|
||||
This is the cause of issue [#2949][issue2949].
|
||||
|
||||
- **Subscriptions are terminated by disconnection.** When the connection to the
|
||||
client is interrupted, the subscription is silently dropped.
|
||||
|
||||
This is a reasonable behavior, but it matters because a client whose
|
||||
subscription is dropped gets no useful error feedback, just a closed
|
||||
connection. Should they try again? Is the node overloaded? Was the client
|
||||
too slow? Did the caller forget to respond to pings? Debugging these kinds
|
||||
of failures is unnecessarily painful.
|
||||
|
||||
Websockets compound this, because websocket connections time out if no
|
||||
traffic is seen for a while, and keeping them alive requires active
|
||||
cooperation between the client and server. With a plain TCP socket, liveness
|
||||
is handled transparently by the keepalive mechanism. On a websocket,
|
||||
however, one side has to occasionally send a PING (if the connection is
|
||||
otherwise idle). The other side must return a matching PONG in time, or the
|
||||
connection is dropped. Apart from being tedious, this is highly susceptible
|
||||
to CPU load.
|
||||
|
||||
The Tendermint Go implementation automatically sends and responds to pings.
|
||||
Clients in other languages (or not wanting to use the Tendermint libraries)
|
||||
need to handle it explicitly. This burdens the client for no practical
|
||||
benefit: A subscriber has no information about when matching events may be
|
||||
available, so it shouldn't have to participate in keeping the connection
|
||||
alive.
|
||||
|
||||
- **Mismatched load profiles.** Most of the RPC service is mainly important for
|
||||
low-volume local use, either by the application the node serves (e.g., the
|
||||
ABCI methods) or by the node operator (e.g., the info methods). Event
|
||||
subscription is important for remote clients, and may represent a much higher
|
||||
volume of traffic.
|
||||
|
||||
This matters because both are using the same JSON-RPC mechanism. For
|
||||
low-volume local use, the ergonomics of JSON-RPC are a good fit: It's easy to
|
||||
issue queries from the command line (e.g., using `curl`) or to write scripts
|
||||
that call the RPC methods to monitor the running node.
|
||||
|
||||
For high-volume remote use, JSON-RPC is not such a good fit: Even leaving
|
||||
aside the non-standard delivery protocol mentioned above, the time and memory
|
||||
cost of encoding event data matters for the stability of the node when there
|
||||
can be potentially hundreds of subscribers. Moreover, a subscription is
|
||||
long-lived compared to most RPC methods, in that it may persist as long the
|
||||
node is active.
|
||||
|
||||
- **Mismatched security profiles.** The RPC service exports several methods
|
||||
that should not be open to arbitrary remote callers, both for correctness
|
||||
reasons (e.g., `remove_tx` and `broadcast_tx_*`) and for operational
|
||||
stability reasons (e.g., `tx_search`). A node may still need to expose
|
||||
events, however, to support UI tools.
|
||||
|
||||
This matters, because all the methods share the same network endpoint. While
|
||||
it is possible to block the top-level GET and POST handlers with a proxy,
|
||||
exposing the `/websocket` handler exposes not _only_ the event subscription
|
||||
methods, but the rest of the service as well.
|
||||
|
||||
### Possible Improvements
|
||||
|
||||
There are several things we could do to improve the experience of developers
|
||||
who need to subscribe to events from the consensus node. These are not all
|
||||
mutually exclusive.
|
||||
|
||||
1. **Split event subscription into a separate service**. Instead of exposing
|
||||
event subscription on the same endpoint as the rest of the RPC service,
|
||||
dedicate a separate endpoint on the node for _only_ event subscription. The
|
||||
rest of the RPC services (_sans_ events) would remain as-is.
|
||||
|
||||
This would make it easy to disable or firewall outside access to sensitive
|
||||
RPC methods, without blocking access to event subscription (and vice versa).
|
||||
This is probably worth doing, even if we don't take any of the other steps
|
||||
described here.
|
||||
|
||||
2. **Use a different protocol for event subscription.** There are various ways
|
||||
we could approach this, depending how much we're willing to shake up the
|
||||
current API. Here are sketches of a few options:
|
||||
|
||||
- Keep the websocket, but rework the API to be more JSON-RPC compliant,
|
||||
perhaps by converting event delivery into notifications. This is less
|
||||
up-front change for existing clients, but retains all of the existing
|
||||
implementation complexity, and doesn't contribute much toward more serious
|
||||
performance and UX improvements later.
|
||||
|
||||
- Switch from websocket to plain HTTP, and rework the subscription API to
|
||||
use a more conventional request/response pattern instead of streaming.
|
||||
This is a little more up-front work for existing clients, but leverages
|
||||
better library support for clients not written in Go.
|
||||
|
||||
The protocol would become more chatty, but we could mitigate that with
|
||||
batching, and in return we would get more control over what to do about
|
||||
slow clients: Instead of simply silently dropping them, as we do now, we
|
||||
could drop messages and signal the client that they missed some data ("M
|
||||
dropped messages since your last poll").
|
||||
|
||||
This option is probably the best balance between work, API change, and
|
||||
benefit, and has a nice incidental effect that it would be easier to debug
|
||||
subscriptions from the command-line, like the other RPC methods.
|
||||
|
||||
- Switch to gRPC: Preserves a persistent connection and gives us a more
|
||||
efficient binary wire format (protobuf), at the cost of much more work for
|
||||
clients and harder debugging. This may be the best option if performance
|
||||
and server load are our top concerns.
|
||||
|
||||
Given that we are currently using JSON-RPC, however, I'm not convinced the
|
||||
costs of encoding and sending messages on the event subscription channel
|
||||
are the limiting factor on subscription efficiency, however.
|
||||
|
||||
3. **Delegate event subscriptions to a proxy.** Give responsibility for
|
||||
managing event subscription to a proxy that runs separately from the node,
|
||||
and switch the node to push events to the proxy (like a webhook) instead of
|
||||
serving subscribers directly. This is more work for the operator (another
|
||||
process to configure and run) but may scale better for big networks.
|
||||
|
||||
I mention this option for completeness, but making this change would be a
|
||||
fairly substantial project. If we want to consider shifting responsibility
|
||||
for event subscription outside the node anyway, we should probably be more
|
||||
systematic about it. For a more principled approach, see point (4) below.
|
||||
|
||||
4. **Move event subscription downstream of indexing.** We are already planning
|
||||
to give applications more control over event indexing. By extension, we
|
||||
might allow the application to also control how events are filtered,
|
||||
queried, and subscribed. Having the application control these concerns,
|
||||
rather than the node, might make life easier for developers building UI and
|
||||
tools for that application.
|
||||
|
||||
This is a much larger change, so I don't think it is likely to be practical
|
||||
in the near-term, but it's worth considering as a broader option. Some of
|
||||
the existing code for filtering and selection could be made more reusable,
|
||||
so applications would not need to reinvent everything.
|
||||
|
||||
|
||||
## References
|
||||
|
||||
- [Tendermint RPC service][rpc-service]
|
||||
- [Tendermint RPC routes][rpc-methods]
|
||||
- [Discussion of the event system][events]
|
||||
- [Discussion about RPC transport options][rpc-transport] (from RFC 002)
|
||||
- [RFC 6455: The websocket protocol][ws]
|
||||
- [JSON-RPC 2.0 Specification](https://www.jsonrpc.org/specification)
|
||||
|
||||
[rpc-service]: https://docs.tendermint.com/master/rpc/
|
||||
[rpc-methods]: https://github.com/tendermint/tendermint/blob/master/internal/rpc/core/routes.go#L12
|
||||
[events]: ./rfc-005-event-system.rst
|
||||
[rpc-transport]: ./rfc-002-ipc-ecosystem.md#rpc-transport
|
||||
[ws]: https://datatracker.ietf.org/doc/html/rfc6455
|
||||
[json-response]: https://www.jsonrpc.org/specification#response_object
|
||||
[json-notify]: https://www.jsonrpc.org/specification#notification
|
||||
[issue2949]: https://github.com/tendermint/tendermint/issues/2949
|
||||
@@ -0,0 +1,140 @@
|
||||
# RFC 007 : Deterministic Proto Byte Serialization
|
||||
|
||||
## Changelog
|
||||
|
||||
- 09-Dec-2021: Initial draft (@williambanfield).
|
||||
|
||||
## Abstract
|
||||
|
||||
This document discusses the issue of stable byte-representation of serialized messages
|
||||
within Tendermint and describes a few possible routes that could be taken to address it.
|
||||
|
||||
## Background
|
||||
|
||||
We use the byte representations of wire-format proto messages to produce
|
||||
and verify hashes of data within the Tendermint codebase as well as for
|
||||
producing and verifying cryptographic signatures over these signed bytes.
|
||||
|
||||
The protocol buffer [encoding spec][proto-spec-encoding] does not guarantee that the byte representation
|
||||
of a protocol buffer message will be the same between two calls to an encoder.
|
||||
While there is a mode to force the encoder to produce the same byte representation
|
||||
of messages within a single binary, these guarantees are not good enough for our
|
||||
use case in Tendermint. We require multiple different versions of a binary running
|
||||
Tendermint to be able to inter-operate. Additionally, we require that multiple different
|
||||
systems written in _different languages_ be able to participate in different aspects
|
||||
of the protocols of Tendermint and be able to verify the integrity of the messages
|
||||
they each produce.
|
||||
|
||||
While this has not yet created a problem that we know of in a running network, we should
|
||||
make sure to provide stronger guarantees around the serialized representation of the messages
|
||||
used within the Tendermint consensus algorithm to prevent any issue from occurring.
|
||||
|
||||
|
||||
## Discussion
|
||||
|
||||
Proto has the following points of variability that can produce non-deterministic byte representation:
|
||||
|
||||
1. Encoding order of fields within a message.
|
||||
|
||||
Proto allows fields to be encoded in any order and even be repeated.
|
||||
|
||||
2. Encoding order of elements of a repeated field.
|
||||
|
||||
`repeated` fields in a proto message can be serialized in any order.
|
||||
|
||||
3. Presence or absence of default values.
|
||||
|
||||
Types in proto have defined default values similar to Go's zero values.
|
||||
Writing or omitting a default value are both legal ways of encoding a wire message.
|
||||
|
||||
4. Serialization of 'unknown' fields.
|
||||
|
||||
Unknown fields can be present when a message is created by a binary with a newer
|
||||
version of the proto that contains fields that the deserializer in a different
|
||||
binary does not yet know about. Deserializers in binaries that do not know about the field
|
||||
will maintain the bytes of the unknown field but not place them into the deserialized structure.
|
||||
|
||||
We have a few options to consider when producing this stable representation.
|
||||
|
||||
### Options for deterministic byte representation
|
||||
|
||||
#### Use only compliant serializers and constrain field usage
|
||||
|
||||
According to [Cosmos-SDK ADR-27][cosmos-sdk-adr-27], when message types obey a simple
|
||||
set of rules, gogoproto produces a consistent byte representation of serialized messages.
|
||||
This seems promising, although more research is needed to guarantee gogoproto always
|
||||
produces a consistent set of bytes on serialized messages. This would solve the problem
|
||||
within Tendermint as written in Go, but would require ensuring that there are similar
|
||||
serializers written in other languages that produce the same output as gogoproto.
|
||||
|
||||
#### Reorder serialized bytes to ensure determinism.
|
||||
|
||||
The serialized form of a proto message can be transformed into a canonical representation
|
||||
by applying simple rules to the serialized bytes. Re-ordering the serialized bytes
|
||||
would allow Tendermint to produce a canonical byte representation without having to
|
||||
simultaneously maintain a custom proto marshaller.
|
||||
|
||||
This could be implemented as a function in many languages that performed the following
|
||||
producing bytes to sign or hashing:
|
||||
|
||||
1. Does not add any of the data from unknown fields into the type to hash.
|
||||
|
||||
Tendermint should not run into a case where it needs to verify the integrity of
|
||||
data with unknown fields for the following reasons:
|
||||
|
||||
The purpose of checking hash equality within Tendermint is to ensure that
|
||||
its local copy of data matches the data that the network agreed on. There should
|
||||
therefore not be a case where a process is checking hash equality using data that it did not expect
|
||||
to receive. What the data represent may be opaque to the process, such as when checking the
|
||||
transactions in a block, _but the process will still have expected to receive this data_,
|
||||
despite not understanding what their internal structure is. It's not clear what it would
|
||||
mean to verify that a block contains data that a process does not know about.
|
||||
|
||||
The same reasoning applies for signature verification within Tendermint. Processes
|
||||
verify that a digital signature signed over a set of bytes by locally reconstructing the
|
||||
data structure that the digital signature signed using the process's local data.
|
||||
|
||||
2. Reordered all message fields to be in tag-sorted order.
|
||||
|
||||
Tag-sorting top-level fields will place all fields of the same tag in a adjacent
|
||||
to eachother within the serialized representation.
|
||||
|
||||
3. Reordered the contents of all `repeated` fields to be in lexicographically sorted order.
|
||||
|
||||
`repeated` fields will appear in a message as having the same tag but will contain different
|
||||
contents. Therefore, lexicographical sorting will produce a stable ordering of
|
||||
fields with the same tag.
|
||||
|
||||
4. Deleted all default values from the byte representation.
|
||||
|
||||
Encoders can include default values or omit them. Most encoders appear to omit them
|
||||
but we may wish to delete them just to be safe.
|
||||
|
||||
5. Recursively performed these operations on any length-delimited subfields.
|
||||
|
||||
Length delimited fields may contain messages, strings, or just bytes. However,
|
||||
it's not possible to know what data is being represented by such a field.
|
||||
A 'string' may happen to have the same structure as an embedded message and we cannot
|
||||
disambiguate. For this reason, we must apply these same rules to all subfields that
|
||||
may contain messages. Because we cannot know if we have totally mangled the interior 'string'
|
||||
or not, this data should never be deserialized or used for anything beyond hashing.
|
||||
|
||||
A **prototype** implementation by @creachadair of this can be found in [the wirepb repo][wire-pb].
|
||||
This could be implemented in multiple languages more simply than ensuring that there are
|
||||
canonical proto serializers that match in each language.
|
||||
|
||||
### Future work
|
||||
|
||||
We should add clear documentation to the Tendermint codebase every time we
|
||||
compare hashes of proto messages or use proto serialized bytes to produces a
|
||||
digital signatures that we have been careful to ensure that the hashes are performed
|
||||
properly.
|
||||
|
||||
### References
|
||||
|
||||
[proto-spec-encoding]: https://developers.google.com/protocol-buffers/docs/encoding
|
||||
[spec-issue]: https://github.com/tendermint/tendermint/issues/5005
|
||||
[cosmos-sdk-adr-27]: https://github.com/cosmos/cosmos-sdk/blob/master/docs/architecture/adr-027-deterministic-protobuf-serialization.md
|
||||
[cer-proto-3]: https://github.com/regen-network/canonical-proto3
|
||||
[wire-pb]: https://github.com/creachadair/wirepb
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
# RFC 008: Don't Panic
|
||||
|
||||
## Changelog
|
||||
|
||||
- 2021-12-17: initial draft (@tychoish)
|
||||
|
||||
## Abstract
|
||||
|
||||
Today, the Tendermint core codebase has panics in a number of cases as
|
||||
a response to exceptional situations. These panics complicate testing,
|
||||
and might make tendermint components difficult to use as a library in
|
||||
some circumstances. This document outlines a project of converting
|
||||
panics to errors and describes the situations where its safe to
|
||||
panic.
|
||||
|
||||
## Background
|
||||
|
||||
Panics in Go are a great mechanism for aborting the current execution
|
||||
for truly exceptional situations (e.g. memory errors, data corruption,
|
||||
processes initialization); however, because they resemble exceptions
|
||||
in other languages, it can be easy to over use them in the
|
||||
implementation of software architectures. This certainly happened in
|
||||
the history of Tendermint, and as we embark on the project of
|
||||
stabilizing the package, we find ourselves in the right moment to
|
||||
reexamine our use of panics, and largely where panics happen in the
|
||||
code base.
|
||||
|
||||
There are still some situations where panics are acceptable and
|
||||
desireable, but it's important that Tendermint, as a project, comes to
|
||||
consensus--perhaps in the text of this document--on the situations
|
||||
where it is acceptable to panic.
|
||||
|
||||
### References
|
||||
|
||||
- [Defer Panic and Recover](https://go.dev/blog/defer-panic-and-recover)
|
||||
- [Why Go gets exceptions right](https://dave.cheney.net/tag/panic)
|
||||
- [Don't panic](https://dave.cheney.net/practical-go/presentations/gophercon-singapore-2019.html#_dont_panic)
|
||||
|
||||
## Discussion
|
||||
|
||||
### Acceptable Panics
|
||||
|
||||
#### Initialization
|
||||
|
||||
It is unambiguously safe (and desireable) to panic in `init()`
|
||||
functions in response to any kind of error. These errors are caught by
|
||||
tests, and occur early enough in process initialization that they
|
||||
won't cause unexpected runtime crashes.
|
||||
|
||||
Other code that is called early in process initialization MAY panic,
|
||||
in some situations if it's not possible to return an error or cause
|
||||
the process to abort early, although these situations should be
|
||||
vanishingly slim.
|
||||
|
||||
#### Data Corruption
|
||||
|
||||
If Tendermint code encounters an inconsistency that could be
|
||||
attributed to data corruption or a logical impossibility it is safer
|
||||
to panic and crash the process than continue to attempt to make
|
||||
progress in these situations.
|
||||
|
||||
Examples including reading data out of the storage engine that
|
||||
is invalid or corrupt, or encountering an ambiguous situation where
|
||||
the process should halt. Generally these forms of corruption are
|
||||
detected after interacting with a trusted but external data source,
|
||||
and reflect situations where the author thinks its safer to terminate
|
||||
the process immediately rather than allow execution to continue.
|
||||
|
||||
#### Unrecoverable Consensus Failure
|
||||
|
||||
In general, a panic should be used in the case of unrecoverable
|
||||
consensus failures. If a process detects that the network is
|
||||
behaving in an incoherent way and it does not have a clearly defined
|
||||
and mechanism for recovering, the process should panic.
|
||||
|
||||
#### Static Validity
|
||||
|
||||
It is acceptable to panic for invariant violations, within a library
|
||||
or package, in situations that should be statically impossible,
|
||||
because there is no way to make these kinds of assertions at compile
|
||||
time.
|
||||
|
||||
For example, type-asserting `interface{}` values returned by
|
||||
`container/list` and `container/heap` (and similar), is acceptable,
|
||||
because package authors should have exclusive control of the inputs to
|
||||
these containers. Packages should not expose the ability to add
|
||||
arbitrary values to these data structures.
|
||||
|
||||
#### Controlled Panics Within Libraries
|
||||
|
||||
In some algorithms with highly recursive structures or very nested
|
||||
call patterns, using a panic, in combination with conditional recovery
|
||||
handlers results in more manageable code. Ultimately this is a limited
|
||||
application, and implementations that use panics internally should
|
||||
only recover conditionally, filtering out panics rather than ignoring
|
||||
or handling all panics.
|
||||
|
||||
#### Request Handling
|
||||
|
||||
Code that handles responses to incoming/external requests
|
||||
(e.g. `http.Handler`) should avoid panics, but practice this isn't
|
||||
totally possible, and it makes sense that request handlers have some
|
||||
kind of default recovery mechanism that will prevent one request from
|
||||
terminating a service.
|
||||
|
||||
### Unacceptable Panics
|
||||
|
||||
In **no** other situation is it acceptable for the code to panic:
|
||||
|
||||
- there should be **no** controlled panics that callers are required
|
||||
to handle across library/package boundaries.
|
||||
- callers of library functions should not expect panics.
|
||||
- ensuring that arbitrary go routines can't panic.
|
||||
- ensuring that there are no arbitrary panics in core production code,
|
||||
espically code that can run at any time during the lifetime of a
|
||||
process.
|
||||
- all test code and fixture should report normal test assertions with
|
||||
a mechanism like testify's `require` assertion rather than calling
|
||||
panic directly.
|
||||
|
||||
The goal of this increased "panic rigor" is to ensure that any escaped
|
||||
panic is reflects a fixable bug in Tendermint.
|
||||
|
||||
### Removing Panics
|
||||
|
||||
The process for removing panics involve a few steps, and will be part
|
||||
of an ongoing process of code modernization:
|
||||
|
||||
- converting existing explicit panics to errors in cases where it's
|
||||
possible to return an error, the errors can and should be handled, and returning
|
||||
an error would not lead to data corruption or cover up data
|
||||
corruption.
|
||||
|
||||
- increase rigor around operations that can cause runtime errors, like
|
||||
type assertions, nil pointer errors, array bounds access issues, and
|
||||
either avoid these situations or return errors where possible.
|
||||
|
||||
- remove generic panic handlers which could cover and hide known
|
||||
panics.
|
||||
@@ -0,0 +1,128 @@
|
||||
# RFC 009 : Consensus Parameter Upgrade Considerations
|
||||
|
||||
## Changelog
|
||||
|
||||
- 06-Jan-2011: Initial draft (@williambanfield).
|
||||
|
||||
## Abstract
|
||||
|
||||
This document discusses the challenges of adding additional consensus parameters
|
||||
to Tendermint and proposes a few solutions that can enable addition of consensus
|
||||
parameters in a backwards-compatible way.
|
||||
|
||||
## Background
|
||||
|
||||
This section provides an overview of the issues of adding consensus parameters
|
||||
to Tendermint.
|
||||
|
||||
### Hash Compatibility
|
||||
|
||||
Tendermint produces a hash of a subset of the consensus parameters. The values
|
||||
that are hashed currently are the `BlockMaxGas` and the `BlockMaxSize`. These
|
||||
are currently in the [HashedParams struct][hashed-params]. This hash is included
|
||||
in the block and validators use it to validate that their local view of the consensus
|
||||
parameters matches what the rest of the network is configured with.
|
||||
|
||||
Any new consensus parameters added to Tendermint should be included in this
|
||||
hash. This presents a challenge for verification of historical blocks when consensus
|
||||
parameters are added. If a network produced blocks with a version of Tendermint that
|
||||
did not yet have the new consensus parameters, the parameter hash it produced will
|
||||
not reference the new parameters. Any nodes joining the network with the newer
|
||||
version of Tendermint will have the new consensus parameters. Tendermint will need
|
||||
to handle this case so that new versions of Tendermint with new consensus parameters
|
||||
can still validate old blocks correctly without having to do anything overly complex
|
||||
or hacky.
|
||||
|
||||
### Allowing Developer-Defined Values and the `EndBlock` Problem
|
||||
|
||||
When new consensus parameters are added, application developers may wish to set
|
||||
values for them so that the developer-defined values may be used as soon as the
|
||||
software upgrades. We do not currently have a clean mechanism for handling this.
|
||||
|
||||
Consensus parameter updates are communicated from the application to Tendermint
|
||||
within `EndBlock` of some height `H` and take effect at the next height, `H+1`.
|
||||
This means that for updates that add a consensus parameter, there is a single
|
||||
height where the new parameters cannot take effect. The parameters did not exist
|
||||
in the version of the software that emitted the `EndBlock` response for height `H-1`,
|
||||
so they cannot take effect at height `H`. The first height that the updated params
|
||||
can take effect is height `H+1`. As of now, height `H` must run with the defaults.
|
||||
|
||||
## Discussion
|
||||
|
||||
### Hash Compatibility
|
||||
|
||||
This section discusses possible solutions to the problem of maintaining backwards-compatibility
|
||||
of hashed parameters while adding new parameters.
|
||||
|
||||
#### Never Hash Defaults
|
||||
|
||||
One solution to the problem of backwards-compatibility is to never include parameters
|
||||
in the hash if the are using the default value. This means that blocks produced
|
||||
before the parameters existed will have implicitly been created with the defaults.
|
||||
This works because any software with newer versions of Tendermint must be using the
|
||||
defaults for new parameters when validating old blocks since the defaults can not
|
||||
have been updated until a height at which the parameters existed.
|
||||
|
||||
#### Only Update HashedParams on Hash-Breaking Releases
|
||||
|
||||
An alternate solution to never hashing defaults is to not update the hashed
|
||||
parameters on non-hash-breaking releases. This means that when new consensus
|
||||
parameters are added to Tendermint, there may be a release that makes use of the
|
||||
parameters but does not verify that they are the same across all validators by
|
||||
referencing them in the hash. This seems reasonably safe given the fact that
|
||||
only a very far subset of the consensus parameters are currently verified at all.
|
||||
|
||||
#### Version The Consensus Parameter Hash Scheme
|
||||
|
||||
The upcoming work on [soft upgrades](https://github.com/tendermint/spec/pull/222)
|
||||
proposes applying different hashing rules depending on the active block version.
|
||||
The consensus parameter hash could be versioned in the same way. When different
|
||||
block versions are used, a different set of consensus parameters will be included
|
||||
in the hash.
|
||||
|
||||
### Developer Defined Values
|
||||
|
||||
This section discusses possible solutions to the problem of allowing application
|
||||
developers to define values for the new parameters during the upgrade that adds
|
||||
the parameters.
|
||||
|
||||
#### Using `InitChain` for New Values
|
||||
|
||||
One solution to the problem of allowing application developers to define values
|
||||
for new consensus parameters is to call the `InitChain` ABCI method on application
|
||||
startup and fetch the value for any new consensus parameters. The [response object][init-chain-response]
|
||||
contains a field for `ConsensusParameter` updates so this may serve as a natural place
|
||||
to put this logic.
|
||||
|
||||
This poses a few difficulties. Nodes replaying old blocks while running new
|
||||
software do not ever call `InitChain` after the initial time. They will therefore
|
||||
not have a way to determine that the parameters changed at some height by using a
|
||||
call to `InitChain`. The `EndBlock` response is how parameter changes at a height
|
||||
are currently communicated to Tendermint and conflating these cases seems risky.
|
||||
|
||||
#### Force Defaults For Single Height
|
||||
|
||||
An alternate option is to not use `InitChain` and instead require chains to use the
|
||||
default values of the new parameters for a single height.
|
||||
|
||||
As documented in the upcoming [ADR-74][adr-74], popular chains often simply use the default
|
||||
values. Additionally, great care is being taken to ensure that logic governed by upcoming
|
||||
consensus parameters is not liveness-breaking. This means that, at worst-case,
|
||||
chains will experience a single slow height while waiting for the new values to
|
||||
by applied.
|
||||
|
||||
#### Add a new `UpgradeChain` method
|
||||
|
||||
An additional method for allowing chains to update the consensus parameters that
|
||||
do not yet exist is to add a new `UpgradeChain` method to `ABCI`. The upgrade chain
|
||||
method would be called when the chain detects that the version of block that it
|
||||
is about to produce does not match the previous block. This method would be called
|
||||
after `EndBlock` and would return the set of consensus parameters to use at the
|
||||
next height. It would therefore give an application the chance to set the new
|
||||
consensus parameters before running a height with these new parameter.
|
||||
|
||||
### References
|
||||
|
||||
[hashed-params]: https://github.com/tendermint/tendermint/blob/0ae974e63911804d4a2007bd8a9b3ad81d6d2a90/types/params.go#L49
|
||||
[init-chain-response]: https://github.com/tendermint/tendermint/blob/0ae974e63911804d4a2007bd8a9b3ad81d6d2a90/abci/types/types.pb.go#L1616
|
||||
[adr-74]: https://github.com/tendermint/tendermint/pull/7503
|
||||
@@ -0,0 +1,145 @@
|
||||
==================================
|
||||
RFC 010: Peer to Peer Light Client
|
||||
==================================
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2022-01-21: Initial draft (@tychoish)
|
||||
|
||||
Abstract
|
||||
--------
|
||||
|
||||
The dependency on access to the RPC system makes running or using the light
|
||||
client more complicated than it should be, because in practice node operators
|
||||
choose to restrict access to these end points (often correctly.) There is no
|
||||
deep dependency for the light client on the RPC system, and there is a
|
||||
persistent notion that "make a p2p light client" is a solution to this
|
||||
operational limitation. This document explores the implications and
|
||||
requirements of implementing a p2p-based light client, as well as the
|
||||
possibilities afforded by this implementation.
|
||||
|
||||
Background
|
||||
----------
|
||||
|
||||
High Level Design
|
||||
~~~~~~~~~~~~~~~~~
|
||||
|
||||
From a high level, the light client P2P implementation, is relatively straight
|
||||
forward, but is orthogonal to the P2P-backed statesync implementation that
|
||||
took place during the 0.35 cycle. The light client only really needs to be
|
||||
able to request (and receive) a `LightBlock` at a given height. To support
|
||||
this, a new Reactor would run on every full node and validator which would be
|
||||
able to service these requests. The workload would be entirely
|
||||
request-response, and the implementation of the reactor would likely be very
|
||||
straight forward, and the implementation of the provider is similarly
|
||||
relatively simple.
|
||||
|
||||
The complexity of the project focuses around peer discovery, handling when
|
||||
peers disconnect from the light clients, and how to change the current P2P
|
||||
code to appropriately handle specialized nodes.
|
||||
|
||||
I believe it's safe to assume that much of the current functionality of the
|
||||
current ``light`` mode would *not* need to be maintained: there is no need to
|
||||
proxy the RPC endpoints over the P2P layer and there may be no need to run a
|
||||
node/process for the p2p light client (e.g. all use of this will be as a
|
||||
client.)
|
||||
|
||||
The ability to run light clients using the RPC system will continue to be
|
||||
maintained.
|
||||
|
||||
LibP2P
|
||||
~~~~~~
|
||||
|
||||
While some aspects of the P2P light client implementation are orthogonal to
|
||||
LibP2P project, it's useful to think about the ways that these efforts may
|
||||
combine or interact.
|
||||
|
||||
We expect to be able to leverage libp2p tools to provide some kind of service
|
||||
discovery for tendermint-based networks. This means that it will be possible
|
||||
for the p2p stack to easily identify specialized nodes, (e.g. light clients)
|
||||
thus obviating many of the design challenges with providing this feature in
|
||||
the context of the current stack.
|
||||
|
||||
Similarly, libp2p makes it possible for a project to be able back their non-Go
|
||||
light clients, without the major task of first implementing Tendermint's p2p
|
||||
connection handling. We should identify if there exist users (e.g. the go IBC
|
||||
relayer, it's maintainers, and operators) who would be able to take advantage
|
||||
of p2p light client, before switching to libp2p. To our knowledge there are
|
||||
limited implementations of this p2p protocol (a simple implementation without
|
||||
secret connection support exists in rust but it has not been used in
|
||||
production), and it seems unlikely that a team would implement this directly
|
||||
ahead of its impending removal.
|
||||
|
||||
User Cases
|
||||
~~~~~~~~~~
|
||||
|
||||
This RFC makes a few assumptions about the use cases and users of light
|
||||
clients in tendermint.
|
||||
|
||||
The most active and delicate use cases for light clients is in the
|
||||
implementation of the IBC relayer. Thus, we expect that providing P2P light
|
||||
clients might increase the reliability of relayers and reduce the cost of
|
||||
running a relayer, because relayer operators won't have to decide between rely
|
||||
on public RPC endpoints (unreliable) or running their own full nodes
|
||||
(expensive.) This also assumes that there are *no* other uses of the RPC in
|
||||
the relayer, and unless the relayers have the option of dropping all RPC use,
|
||||
it's unclear if a P2P light client will actually be able to successfully
|
||||
remove the dependency on the RPC system.
|
||||
|
||||
Given that the primary relayer implementation is Hermes (rust,) it might be
|
||||
safe to deliver a version of Tendermint that adds a light client rector in
|
||||
the full nodes, but that does not provide an implementation of a Go light
|
||||
client. This either means that the rust implementation would need support for
|
||||
the legacy P2P connection protocol or wait for the libp2p implementation.
|
||||
|
||||
Client side light client (e.g. wallets, etc.) users may always want to use (a
|
||||
subset) of the RPC rather than connect to the P2P network for an ephemeral
|
||||
use.
|
||||
|
||||
Discussion
|
||||
----------
|
||||
|
||||
Implementation Questions
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Most of the complication in the is how to have a long lived light client node
|
||||
that *only* runs the light client reactor, as this raises a few questions:
|
||||
|
||||
- would users specify a single P2P node to connect to when creating a light
|
||||
client or would they also need/want to discover peers?
|
||||
|
||||
- **answer**: most light client use cases won't care much about selecting
|
||||
peers (and those that do can either disable PEX and specify persistent
|
||||
peers, *or* use the RPC light client.)
|
||||
|
||||
- how do we prevent full nodes and validators from allowing their peer slots,
|
||||
which are typically limited, from filling with light clients? If
|
||||
light-clients aren't limited, how do we prevent light clients from consuming
|
||||
resources on consensus nodes?
|
||||
|
||||
- **answer**: I think we can institute an internal cap on number of light
|
||||
client connections to accept and also elide light client nodes from PEX
|
||||
(pre-libp2p, if we implement this.) I believe that libp2p should provide
|
||||
us with the kind of service discovery semantics for network connectivity
|
||||
that would obviate this issue.
|
||||
|
||||
- when a light client disconnects from its peers will it need to reset its
|
||||
internal state (cache)? does this change if it connects to the same peers?
|
||||
|
||||
- **answer**: no, the internal state only needs to be reset if the light
|
||||
client detects an invalid block or other divergence, and changing
|
||||
witnesses--which will be more common with a p2p light client--need not
|
||||
invalidate the cache.
|
||||
|
||||
These issues are primarily present given that the current peer management later
|
||||
does not have a particularly good service discovery mechanism nor does it have
|
||||
a very sophisticated way of identifying nodes of different types or modes.
|
||||
|
||||
Report Evidence
|
||||
~~~~~~~~~~~~~~~
|
||||
|
||||
The current light client implementation currently has the ability to report
|
||||
observed evidence. Either the notional light client reactor needs to be able
|
||||
to handle these kinds of requests *or* all light client nodes need to also run
|
||||
the evidence reactor. This could be configured at runtime.
|
||||
@@ -0,0 +1,162 @@
|
||||
# RFC 011: Remove Gas From Tendermint
|
||||
|
||||
## Changelog
|
||||
|
||||
- 03-Feb-2022: Initial draft (@williambanfield).
|
||||
- 10-Feb-2022: Update in response to feedback (@williambanfield).
|
||||
- 11-Feb-2022: Add reflection on MaxGas during consensus (@williambanfield).
|
||||
|
||||
## Abstract
|
||||
|
||||
In the v0.25.0 release, Tendermint added a mechanism for tracking 'Gas' in the mempool.
|
||||
At a high level, Gas allows applications to specify how much it will cost the network,
|
||||
often in compute resources, to execute a given transaction. While such a mechanism is common
|
||||
in blockchain applications, it is not generalizable enough to be a maintained as a part
|
||||
of Tendermint. This RFC explores the possibility of removing the concept of Gas from
|
||||
Tendermint while still allowing applications the power to control the contents of
|
||||
blocks to achieve similar goals.
|
||||
|
||||
## Background
|
||||
|
||||
The notion of Gas was included in the original Ethereum whitepaper and exists as
|
||||
an important feature of the Ethereum blockchain.
|
||||
|
||||
The [whitepaper describes Gas][eth-whitepaper-messages] as an Anti-DoS mechanism. The Ethereum Virtual Machine
|
||||
provides a Turing complete execution platform. Without any limitations, malicious
|
||||
actors could waste computation resources by directing the EVM to perform large
|
||||
or even infinite computations. Gas serves as a metering mechanism to prevent this.
|
||||
|
||||
Gas appears to have been added to Tendermint multiple times, initially as part of
|
||||
a now defunct `/vm` package, and in its most recent iteration [as part of v0.25.0][gas-add-pr]
|
||||
as a mechanism to limit the transactions that will be included in the block by an additional
|
||||
parameter.
|
||||
|
||||
Gas has gained adoption within the Cosmos ecosystem [as part of the Cosmos SDK][cosmos-sdk-gas].
|
||||
The SDK provides facilities for tracking how much 'Gas' a transaction is expected to take
|
||||
and a mechanism for tracking how much gas a transaction has already taken.
|
||||
|
||||
Non-SDK applications also make use of the concept of Gas. Anoma appears to implement
|
||||
[a gas system][anoma-gas] to meter the transactions it executes.
|
||||
|
||||
While the notion of gas is present in projects that make use of Tendermint, it is
|
||||
not a concern of Tendermint's. Tendermint's value and goal is producing blocks
|
||||
via a distributed consensus algorithm. Tendermint relies on the application specific
|
||||
code to decide how to handle the transactions Tendermint has produced (or if the
|
||||
application wants to consider them at all). Gas is an application concern.
|
||||
|
||||
Our implementation of Gas is not currently enforced by consensus. Our current validation check that
|
||||
occurs during block propagation does not verify that the block is under the configured `MaxGas`.
|
||||
Ensuring that the transactions in a proposed block do not exceed `MaxGas` would require
|
||||
input from the application during propagation. The `ProcessProposal` method introduced
|
||||
as part of ABCI++ would enable such input but would further entwine Tendermint and
|
||||
the application. The issue of checking `MaxGas` during block propagation is important
|
||||
because it demonstrates that the feature as it currently exists is not implemented
|
||||
as fully as it perhaps should be.
|
||||
|
||||
Our implementation of Gas is causing issues for node operators and relayers. At
|
||||
the moment, transactions that overflow the configured 'MaxGas' can be silently rejected
|
||||
from the mempool. Overflowing MaxGas is the _only_ way that a transaction can be considered
|
||||
invalid that is not directly a result of failing the `CheckTx`. Operators, and the application,
|
||||
do not know that a transaction was removed from the mempool for this reason. A stateless check
|
||||
of this nature is exactly what `CheckTx` exists for and there is no reason for the mempool
|
||||
to keep track of this data separately. A special [MempoolError][add-mempool-error] field
|
||||
was added in v0.35 to communicate to clients that a transaction failed after `CheckTx`.
|
||||
While this should alleviate the pain for operators wishing to understand if their
|
||||
transaction was included in the mempool, it highlights that the abstraction of
|
||||
what is included in the mempool is not currently well defined.
|
||||
|
||||
Removing Gas from Tendermint and the mempool would allow for the mempool to be a better
|
||||
abstraction: any transaction that arrived at `CheckTx` and passed the check will either be
|
||||
a candidate for a later block or evicted after a TTL is reached or to make room for
|
||||
other, higher priority transactions. All other transactions are completely invalid and can be discarded forever.
|
||||
|
||||
Removing gas will not be completely straightforward. It will mean ensuring that
|
||||
equivalent functionality can be implemented outside of the mempool using the mempool's API.
|
||||
|
||||
## Discussion
|
||||
|
||||
This section catalogs the functionality that will need to exist within the Tendermint
|
||||
mempool to allow Gas to be removed and replaced by application-side bookkeeping.
|
||||
|
||||
### Requirement: Provide Mempool Tx Sorting Mechanism
|
||||
|
||||
Gas produces a market for inclusion in a block. On many networks, a [gas fee][cosmos-sdk-fees] is
|
||||
included in pending transactions. This fee indicates how much a user is willing to
|
||||
pay per unit of execution and the fees are distributed to validators.
|
||||
|
||||
Validators wishing to extract higher gas fees are incentivized to include transactions
|
||||
with the highest listed gas fees into each block. This produces a natural ordering
|
||||
of the pending transactions. Applications wishing to implement a gas mechanism need
|
||||
to be able to order the transactions in the mempool. This can trivially be accomplished
|
||||
by sorting transactions using the `priority` field available to applications as part of
|
||||
v0.35's `ResponseCheckTx` message.
|
||||
|
||||
### Requirement: Allow Application-Defined Block Resizing
|
||||
|
||||
When creating a block proposal, Tendermint pulls a set of possible transactions out of
|
||||
the mempool to include in the next block. Tendermint uses MaxGas to limit the set of transactions
|
||||
it pulls out of the mempool fetching a set of transactions whose sum is less than MaxGas.
|
||||
|
||||
By removing gas tracking from Tendermint's mempool, Tendermint will need to provide a way for
|
||||
applications to determine an acceptable set of transactions to include in the block.
|
||||
|
||||
This is what the new ABCI++ `PrepareProposal` method is useful for. Applications
|
||||
that wish to limit the contents of a block by an application-defined limit may
|
||||
do so by removing transactions from the proposal it is passed during `PrepareProposal`.
|
||||
Applications wishing to reach parity with the current Gas implementation may do
|
||||
so by creating an application-side limit: filtering out transactions from
|
||||
`PrepareProposal` the cause the proposal the exceed the maximum gas. Additionally,
|
||||
applications can currently opt to have all transactions in the mempool delivered
|
||||
during `PrepareProposal` by passing `-1` for `MaxGas` and `MaxBytes` into
|
||||
[ReapMaxBytesMaxGas][reap-max-bytes-max-gas].
|
||||
|
||||
### Requirement: Handle Transaction Metadata
|
||||
|
||||
Moving the gas mechanism into applications adds an additional piece of complexity
|
||||
to applications. The application must now track how much gas it expects a transaction
|
||||
to consume. The mempool currently handles this bookkeeping responsibility and uses the estimated
|
||||
gas to determine the set of transactions to include in the block. In order to task
|
||||
the application with keeping track of this metadata, we should make it easier for the
|
||||
application to do so. In general, we'll want to keep only one copy of this type
|
||||
of metadata in the program at a time, either in the application or in Tendermint.
|
||||
|
||||
The following sections are possible solutions to the problem of storing transaction
|
||||
metadata without duplication.
|
||||
|
||||
#### Metadata Handling: EvictTx Callback
|
||||
|
||||
A possible approach to handling transaction metadata is by adding a new `EvictTx`
|
||||
ABCI method. Whenever the mempool is removing a transaction, either because it has
|
||||
reached its TTL or because it failed `RecheckTx`, `EvictTx` would be called with
|
||||
the transaction hash. This would indicate to the application that it could free any
|
||||
metadata it was storing about the transaction such as the computed gas fee.
|
||||
|
||||
Eviction callbacks are pretty common in caching systems, so this would be very
|
||||
well-worn territory.
|
||||
|
||||
#### Metadata Handling: Application-Specific Metadata Field(s)
|
||||
|
||||
An alternative approach to handling transaction metadata would be would be the
|
||||
addition of a new application-metadata field in the `ResponseCheckTx`. This field
|
||||
would be a protocol buffer message whose contents were entirely opaque to Tendermint.
|
||||
The application would be responsible for marshalling and unmarshalling whatever data
|
||||
it stored in this field. During `PrepareProposal`, the application would be passed
|
||||
this metadata along with the transaction, allowing the application to use it to perform
|
||||
any necessary filtering.
|
||||
|
||||
If either of these proposed metadata handling techniques are selected, it's likely
|
||||
useful to enable applications to gossip metadata along with the transaction it is
|
||||
gossiping. This could easily take the form of an opaque proto message that is
|
||||
gossiped along with the transaction.
|
||||
|
||||
## References
|
||||
|
||||
[eth-whitepaper-messages]: https://ethereum.org/en/whitepaper/#messages-and-transactions
|
||||
[gas-add-pr]: https://github.com/tendermint/tendermint/pull/2360
|
||||
[cosmos-sdk-gas]: https://github.com/cosmos/cosmos-sdk/blob/c00cedb1427240a730d6eb2be6f7cb01f43869d3/docs/basics/gas-fees.md
|
||||
[cosmos-sdk-fees]: https://github.com/cosmos/cosmos-sdk/blob/c00cedb1427240a730d6eb2be6f7cb01f43869d3/docs/basics/tx-lifecycle.md#gas-and-fees
|
||||
[anoma-gas]: https://github.com/anoma/anoma/blob/6974fe1532a59db3574fc02e7f7e65d1216c1eb2/docs/src/specs/ledger.md#transaction-execution
|
||||
[cosmos-sdk-fee]: https://github.com/cosmos/cosmos-sdk/blob/c00cedb1427240a730d6eb2be6f7cb01f43869d3/types/tx/tx.pb.go#L780-L794
|
||||
[issue-7750]: https://github.com/tendermint/tendermint/issues/7750
|
||||
[reap-max-bytes-max-gas]: https://github.com/tendermint/tendermint/blob/1ac58469f32a98f1c0e2905ca1773d9eac7b7103/internal/mempool/types.go#L45
|
||||
[add-mempool-error]: https://github.com/tendermint/tendermint/blob/205bfca66f6da1b2dded381efb9ad3792f9404cf/rpc/coretypes/responses.go#L239
|
||||
@@ -0,0 +1,352 @@
|
||||
# RFC 012: Event Indexing Revisited
|
||||
|
||||
## Changelog
|
||||
|
||||
- 11-Feb-2022: Add terminological notes.
|
||||
- 10-Feb-2022: Updated from review feedback.
|
||||
- 07-Feb-2022: Initial draft (@creachadair)
|
||||
|
||||
## Abstract
|
||||
|
||||
A Tendermint node allows ABCI events associated with block and transaction
|
||||
processing to be "indexed" into persistent storage. The original Tendermint
|
||||
implementation provided a fixed, built-in [proprietary indexer][kv-index] for
|
||||
such events.
|
||||
|
||||
In response to user requests to customize indexing, [ADR 065][adr065]
|
||||
introduced an "event sink" interface that allows developers (at least in
|
||||
theory) to plug in alternative index storage.
|
||||
|
||||
Although ADR-065 was a good first step toward customization, its implementation
|
||||
model does not satisfy all the user requirements. Moreover, this approach
|
||||
leaves some existing technical issues with indexing unsolved.
|
||||
|
||||
This RFC documents these concerns, and discusses some potential approaches to
|
||||
solving them. This RFC does _not_ propose a specific technical decision. It is
|
||||
meant to unify and focus some of the disparate discussions of the topic.
|
||||
|
||||
|
||||
## Background
|
||||
|
||||
We begin with some important terminological context. The term "event" in
|
||||
Tendermint can be confusing, as the same word is used for multiple related but
|
||||
distinct concepts:
|
||||
|
||||
1. **ABCI Events** refer to the key-value metadata attached to blocks and
|
||||
transactions by the application. These values are represented by the ABCI
|
||||
`Event` protobuf message type.
|
||||
|
||||
2. **Consensus Events** refer to the data published by the Tendermint node to
|
||||
its pubsub bus in response to various consensus state transitions and other
|
||||
important activities, such as round updates, votes, transaction delivery,
|
||||
and block completion.
|
||||
|
||||
This confusion is compounded because some "consensus event" values also have
|
||||
"ABCI event" metadata attached to them. Notably, block and transaction items
|
||||
typically have ABCI metadata assigned by the application.
|
||||
|
||||
Indexers and RPC clients subscribed to the pubsub bus receive **consensus
|
||||
events**, but they identify which ones to care about using query expressions
|
||||
that match against the **ABCI events** associated with them.
|
||||
|
||||
In the discussion that follows, we will use the term **event item** to refer to
|
||||
a datum published to or received from the pubsub bus, and **ABCI event** or
|
||||
**event metadata** to refer to the key/value annotations.
|
||||
|
||||
**Indexing** in this context means recording the association between certain
|
||||
ABCI metadata and the blocks or transactions they're attached to. The ABCI
|
||||
metadata typically carry application-specific details like sender and recipient
|
||||
addresses, catgory tags, and so forth, that are not part of consensus but are
|
||||
used by UI tools to find and display transactions of interest.
|
||||
|
||||
The consensus node records the blocks and transactions as part of its block
|
||||
store, but does not persist the application metadata. Metadata persistence is
|
||||
the task of the indexer, which can be (optionally) enabled by the node
|
||||
operator.
|
||||
|
||||
### History
|
||||
|
||||
The [original indexer][kv-index] built in to Tendermint stored index data in an
|
||||
embedded [`tm-db` database][tmdb] with a proprietary key layout.
|
||||
In [ADR 065][adr065], we noted that this implementation has both performance
|
||||
and scaling problems under load. Moreover, the only practical way to query the
|
||||
index data is via the [query filter language][query] used for event
|
||||
subscription. [Issue #1161][i1161] appears to be a motivational context for that ADR.
|
||||
|
||||
To mitigate both of these concerns, we introduced the [`EventSink`][esink]
|
||||
interface, combining the original transaction and block indexer interfaces
|
||||
along with some service plumbing. Using this interface, a developer can plug
|
||||
in an indexer that uses a more efficient storage engine, and provides a more
|
||||
expressive query language. As a proof-of-concept, we built a [PostgreSQL event
|
||||
sink][psql] that exports data to a [PostgreSQL database][postgres].
|
||||
|
||||
Although this approach addressed some of the immediate concerns, there are
|
||||
several issues for custom indexing that have not been fully addressed. Here we
|
||||
will discuss them in more detail.
|
||||
|
||||
For further context, including links to user reports and related work, see also
|
||||
the [Pluggable custom event indexing tracking issue][i7135] issue.
|
||||
|
||||
### Issue 1: Tight Coupling
|
||||
|
||||
The `EventSink` interface supports multiple implementations, but plugging in
|
||||
implementations still requires tight integration with the node. In particular:
|
||||
|
||||
- Any custom indexer must either be written in Go and compiled in to the
|
||||
Tendermint binary, or the developer must write a Go shim to communicate with
|
||||
the implementation and build that into the Tendermint binary.
|
||||
|
||||
- This means to support a custom indexer, it either has to be integrated into
|
||||
the Tendermint core repository, or every installation that uses that indexer
|
||||
must fetch or build a patched version of Tendermint.
|
||||
|
||||
The problem with integrating indexers into Tendermint Core is that every user
|
||||
of Tendermint Core takes a dependency on all supported indexers, including
|
||||
those they never use. Even if the unused code is disabled with build tags,
|
||||
users have to remember to do this or potentially be exposed to security issues
|
||||
that may arise in any of the custom indexers. This is a risk for Tendermint,
|
||||
which is a trust-critical component of all applications built on it.
|
||||
|
||||
The problem with _not_ integrating indexers into Tendermint Core is that any
|
||||
developer who wants to use a particular indexer must now fetch or build a
|
||||
patched version of the core code that includes the custom indexer. Besides
|
||||
being inconvenient, this makes it harder for users to upgrade their node, since
|
||||
they need to either re-apply their patches directly or wait for an intermediary
|
||||
to do it for them.
|
||||
|
||||
Even for developers who have written their applications in Go and link with the
|
||||
consensus node directly (e.g., using the [Cosmos SDK][sdk]), these issues add a
|
||||
potentially significant complication to the build process.
|
||||
|
||||
### Issue 2: Legacy Compatibility
|
||||
|
||||
The `EventSink` interface retains several limitations of the original
|
||||
proprietary indexer. These include:
|
||||
|
||||
- The indexer has no control over which event items are reported. Only the
|
||||
exact block and transaction events that were reported to the original indexer
|
||||
are reported to a custom indexer.
|
||||
|
||||
- The interface requires the implementation to define methods for the legacy
|
||||
search and query API. This requirement comes from the integation with the
|
||||
[event subscription RPC API][event-rpc], but actually supporting these
|
||||
methods is not trivial.
|
||||
|
||||
At present, only the original KV indexer implements the query methods. Even the
|
||||
proof-of-concept PostgreSQL implementation simply reports errors for all calls
|
||||
to these methods.
|
||||
|
||||
Even for a plugin written in Go, implementing these methods "correctly" would
|
||||
require parsing and translating the custom query language over whatever storage
|
||||
platform the indexer uses.
|
||||
|
||||
For a plugin _not_ written in Go, even beyond the cost of integration the
|
||||
developer would have to re-implement the entire query language.
|
||||
|
||||
### Issue 3: Indexing Delays Consensus
|
||||
|
||||
Within the node, indexing hooks in to the same internal pubsub dispatcher that
|
||||
is used to export event items to the [event subscription RPC API][event-rpc].
|
||||
In contrast with RPC subscribers, however, indexing is a "privileged"
|
||||
subscriber: If an RPC subscriber is "too slow", the node may terminate the
|
||||
subscription and disconnect the client. That means that RPC subscribers may
|
||||
lose (miss) event items. The indexer, however, is "unbuffered", and the
|
||||
publisher will never drop or disconnect from it. If the indexer is slow, the
|
||||
publisher will block until it returns, to ensure that no event items are lost.
|
||||
|
||||
In practice, this means that the performance of the indexer has a direct effect
|
||||
on the performance of the consensus node: If the indexer is slow or stalls, it
|
||||
will slow or halt the progress of consensus. Users have already reported this
|
||||
problem even with the built-in indexer (see, for example, [#7247][i7247]).
|
||||
Extending this concern to arbitrary user-defined custom indexers gives that
|
||||
risk a much larger surface area.
|
||||
|
||||
|
||||
## Discussion
|
||||
|
||||
It is not possible to simultaneously guarantee that publishing event items will
|
||||
not delay consensus, and also that all event items of interest are always
|
||||
completely indexed.
|
||||
|
||||
Therefore, our choice is between eliminating delay (and minimizing loss) or
|
||||
eliminating loss (and minimizing delay). Currently, we take the second
|
||||
approach, which has led to user complaints about consensus delays due to
|
||||
indexing and subscription overhead.
|
||||
|
||||
- If we agree that consensus performance supersedes index completeness, our
|
||||
design choices are to constrain the likelihood and frequency of missing event
|
||||
items.
|
||||
|
||||
- If we decide that consensus performance is more important than index
|
||||
completeness, our option is to minimize overhead on the event delivery path
|
||||
and document that indexer plugins constrain the rate of consensus.
|
||||
|
||||
Since we have user reports requesting both properties, we have to choose one or
|
||||
the other. Since the primary job of the consensus engine is to correctly,
|
||||
robustly, reliablly, and efficiently replicate application state across the
|
||||
network, I believe the correct choice is to favor consensus performance.
|
||||
|
||||
An important consideration for this decision is that a node does not index
|
||||
application metadata separately: If indexing is disabled, there is no built-in
|
||||
mechanism to go back and replay or reconstruct the data that an indexer would
|
||||
have stored. The node _does_ store the blockchain itself (i.e., the blocks and
|
||||
their transactions), so potentially some use cases currently handled by the
|
||||
indexer could be handled by the node. For example, allowing clients to ask
|
||||
whether a given transaction ID has been committed to a block could in principle
|
||||
be done without an indexer, since it does not depend on application metadata.
|
||||
|
||||
Inevitably, a question will arise whether we could implement both strategies
|
||||
and toggle between them with a flag. That would be a worst-case scenario,
|
||||
requiring us to maintain the complexity of two very-different operational
|
||||
concerns. If our goal is that Tendermint should be as simple, efficient, and
|
||||
trustworthy as posible, there is not a strong case for making these options
|
||||
configurable: We should pick a side and commit to it.
|
||||
|
||||
### Design Principles
|
||||
|
||||
Although there is no unique "best" solution to the issues described above,
|
||||
there are some specific principles that a solution should include:
|
||||
|
||||
1. **A custom indexer should not require integration into Tendermint core.** A
|
||||
developer or node operator can create, build, deploy, and use a custom
|
||||
indexer with a stock build of the Tendermint consensus node.
|
||||
|
||||
2. **Custom indexers cannot stall consensus.** An indexer that is slow or
|
||||
stalls cannot slow down or prevent core consensus from making progress.
|
||||
|
||||
The plugin interface must give node operators control over the tolerances
|
||||
for acceptable indexer performance, and the means to detect when indexers
|
||||
are falling outside those tolerances, but indexer failures should "fail
|
||||
safe" with respect to consensus (even if that means the indexer may miss
|
||||
some data, in sufficiently-extreme circumstances).
|
||||
|
||||
3. **Custom indexers control which event items they index.** A custom indexer
|
||||
is not limited to only the current transaction and block events, but can
|
||||
observe any event item published by the node.
|
||||
|
||||
4. **Custom indexing is forward-compatible.** Adding new event item types or
|
||||
metadata to the consensus node should not require existing custom indexers
|
||||
to be rebuilt or modified, unless they want to take advantage of the new
|
||||
data.
|
||||
|
||||
5. **Indexers are responsible for answering queries.** An indexer plugin is not
|
||||
required to support the legacy query filter language, nor to be compatible
|
||||
with the legacy RPC endpoints for accessing them. Any APIs for clients to
|
||||
query a custom index are the responsibility of the indexer, not the node.
|
||||
|
||||
### Open Questions
|
||||
|
||||
Given the constraints outlined above, there are important design questions we
|
||||
must answer to guide any specific changes:
|
||||
|
||||
1. **What is an acceptable probability that, given sufficiently extreme
|
||||
operational issues, an indexer might miss some number of events?**
|
||||
|
||||
There are two parts to this question: One is what constitutes an extreme
|
||||
operational problem, the other is how likely we are to miss some number of
|
||||
events items.
|
||||
|
||||
- If the consensus is that no event item must ever be missed, no matter how
|
||||
bad the operational circumstances, then we _must_ accept that indexing can
|
||||
slow or halt consensus arbitrarily. It is impossible to guarantee complete
|
||||
index coverage without potentially unbounded delays.
|
||||
|
||||
- Otherwise, how much data can we afford to lose and how often? For example,
|
||||
if we can ensure no event item will be lost unless the indexer halts for
|
||||
at least five minutes, is that acceptable? What probabilities and time
|
||||
ranges are reasonable for real production environments?
|
||||
|
||||
2. **What level of operational overhead is acceptable to impose on node
|
||||
operators to support indexing?**
|
||||
|
||||
Are node operators willing to configure and run custom indexers as sidecar
|
||||
type processes alongside a node? How much indexer setup above and beyond the
|
||||
work of setting up the underlying node in isolation is tractable in
|
||||
production networks?
|
||||
|
||||
The answer to this question also informs the question of whether we should
|
||||
keep an "in-process" indexing option, and to what extent that option needs
|
||||
to satisfy the suggested design principles.
|
||||
|
||||
Relatedly, to what extent do we need to be concerned about the cost of
|
||||
encoding and sending event items to an external process (e.g., as JSON blobs
|
||||
or protobuf wire messages)? Given that the node already encodes event items
|
||||
as JSON for subscription purposes, the overhead would be negligible for the
|
||||
node itself, but the indexer would have to decode to process the results.
|
||||
|
||||
3. **What (if any) query APIs does the consensus node need to export,
|
||||
independent of the indexer implementation?**
|
||||
|
||||
One typical example is whether the node should be able to answer queries
|
||||
like "is this transaction ID in a block?" Currently, a node cannot answer
|
||||
this query _unless_ it runs the built-in KV indexer. Does the node need to
|
||||
continue to support that query even for nodes that disable the KV indexer,
|
||||
or which use a custom indexer?
|
||||
|
||||
### Informal Design Intent
|
||||
|
||||
The design principles described above implicate several components of the
|
||||
Tendermint node, beyond just the indexer. In the context of [ADR 075][adr075],
|
||||
we are re-working the RPC event subscription API to improve some of the UX
|
||||
issues discussed above for RPC clients. It is our expectation that a solution
|
||||
for pluggable custom indexing will take advantage of some of the same work.
|
||||
|
||||
On that basis, the design approach I am considering for custom indexing looks
|
||||
something like this (subject to refinement):
|
||||
|
||||
1. A custom indexer runs as a separate process from the node.
|
||||
|
||||
2. The indexer subscribes to event items via the ADR 075 events API.
|
||||
|
||||
This means indexers would receive event payloads as JSON rather than
|
||||
protobuf, but since we already have to support JSON encoding for the RPC
|
||||
interface anyway, that should not increase complexity for the node.
|
||||
|
||||
3. The existing PostgreSQL indexer gets reworked to have this form, and no
|
||||
longer built as part of the Tendermint core binary.
|
||||
|
||||
We can retain the code in the core repository as a proof-of-concept, or
|
||||
perhaps create a separate repository with contributed indexers and move it
|
||||
there.
|
||||
|
||||
4. (Possibly) Deprecate and remove the legacy KV indexer, or disable it by
|
||||
default. If we decide to remove it, we can also remove the legacy RPC
|
||||
endpoints for querying the KV indexer.
|
||||
|
||||
If we plan to do this, we should also investigate providing a way for
|
||||
clients to query whether a given transaction ID has landed in a block. That
|
||||
serves a common need, and currently _only_ works if the KV indexer is
|
||||
enabled, but could be addressed more simply using the other data a node
|
||||
already has stored, without having to answer more general queries.
|
||||
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 065: Custom Event Indexing][adr065]
|
||||
- [ADR 075: RPC Event Subscription Interface][adr075]
|
||||
- [Cosmos SDK][sdk]
|
||||
- [Event subscription RPC][event-rpc]
|
||||
- [KV transaction indexer][kv-index]
|
||||
- [Pluggable custom event indexing][i7135] (#7135)
|
||||
- [PostgreSQL event sink][psql]
|
||||
- [PostgreSQL database][postgres]
|
||||
- [Query filter language][query]
|
||||
- [Stream events to postgres for indexing][i1161] (#1161)
|
||||
- [Unbuffered event subscription slow down the consensus][i7247] (#7247)
|
||||
- [`EventSink` interface][esink]
|
||||
- [`tm-db` library][tmdb]
|
||||
|
||||
[adr065]: https://github.com/tendermint/tendermint/blob/master/docs/architecture/adr-065-custom-event-indexing.md
|
||||
[adr075]: https://github.com/tendermint/tendermint/blob/master/docs/architecture/adr-075-rpc-subscription.md
|
||||
[esink]: https://pkg.go.dev/github.com/tendermint/tendermint/internal/state/indexer#EventSink
|
||||
[event-rpc]: https://docs.tendermint.com/master/rpc/#/Websocket/subscribe
|
||||
[i1161]: https://github.com/tendermint/tendermint/issues/1161
|
||||
[i7135]: https://github.com/tendermint/tendermint/issues/7135
|
||||
[i7247]: https://github.com/tendermint/tendermint/issues/7247
|
||||
[kv-index]: https://github.com/tendermint/tendermint/blob/master/internal/state/indexer/tx/kv
|
||||
[postgres]: https://postgresql.org/
|
||||
[psql]: https://github.com/tendermint/tendermint/blob/master/internal/state/indexer/sink/psql
|
||||
[psql]: https://github.com/tendermint/tendermint/blob/master/internal/state/indexer/sink/psql
|
||||
[query]: https://pkg.go.dev/github.com/tendermint/tendermint/internal/pubsub/query/syntax
|
||||
[sdk]: https://github.com/cosmos/cosmos-sdk
|
||||
[tmdb]: https://pkg.go.dev/github.com/tendermint/tm-db#DB
|
||||
@@ -0,0 +1,253 @@
|
||||
# RFC 013: ABCI++
|
||||
|
||||
## Changelog
|
||||
|
||||
- 2020-01-11: initialized
|
||||
- 2022-02-11: Migrate RFC to tendermint repo (Originally [RFC 004](https://github.com/tendermint/spec/pull/254))
|
||||
|
||||
## Author(s)
|
||||
|
||||
- Dev (@valardragon)
|
||||
- Sunny (@sunnya97)
|
||||
|
||||
## Context
|
||||
|
||||
ABCI is the interface between the consensus engine and the application.
|
||||
It defines when the application can talk to consensus during the execution of a blockchain.
|
||||
At the moment, the application can only act at one phase in consensus, immediately after a block has been finalized.
|
||||
|
||||
This restriction on the application prohibits numerous features for the application, including many scalability improvements that are now better understood than when ABCI was first written.
|
||||
For example, many of the scalability proposals can be boiled down to "Make the miner / block proposers / validators do work, so the network does not have to".
|
||||
This includes optimizations such as tx-level signature aggregation, state transition proofs, etc.
|
||||
Furthermore, many new security properties cannot be achieved in the current paradigm, as the application cannot enforce validators do more than just finalize txs.
|
||||
This includes features such as threshold cryptography, and guaranteed IBC connection attempts.
|
||||
We propose introducing three new phases to ABCI to enable these new features, and renaming the existing methods for block execution.
|
||||
|
||||
#### Prepare Proposal phase
|
||||
|
||||
This phase aims to allow the block proposer to perform more computation, to reduce load on all other full nodes, and light clients in the network.
|
||||
It is intended to enable features such as batch optimizations on the transaction data (e.g. signature aggregation, zk rollup style validity proofs, etc.), enabling stateless blockchains with validator provided authentication paths, etc.
|
||||
|
||||
This new phase will only be executed by the block proposer. The application will take in the block header and raw transaction data output by the consensus engine's mempool. It will then return block data that is prepared for gossip on the network, and additional fields to include into the block header.
|
||||
|
||||
#### Process Proposal Phase
|
||||
|
||||
This phase aims to allow applications to determine validity of a new block proposal, and execute computation on the block data, prior to the blocks finalization.
|
||||
It is intended to enable applications to reject block proposals with invalid data, and to enable alternate pipelined execution models. (Such as Ethereum-style immediate execution)
|
||||
|
||||
This phase will be executed by all full nodes upon receiving a block, though on the application side it can do more work in the even that the current node is a validator.
|
||||
|
||||
#### Vote Extension Phase
|
||||
|
||||
This phase aims to allow applications to require their validators do more than just validate blocks.
|
||||
Example usecases of this include validator determined price oracles, validator guaranteed IBC connection attempts, and validator based threshold crypto.
|
||||
|
||||
This adds an app-determined data field that every validator must include with their vote, and these will thus appear in the header.
|
||||
|
||||
#### Rename {BeginBlock, [DeliverTx], EndBlock} to FinalizeBlock
|
||||
|
||||
The prior phases gives the application more flexibility in their execution model for a block, and they obsolete the current methods for how the consensus engine relates the block data to the state machine. Thus we refactor the existing methods to better reflect what is happening in the new ABCI model.
|
||||
|
||||
This rename doesn't on its own enable anything new, but instead improves naming to clarify the expectations from the application in this new communication model. The existing ABCI methods `BeginBlock, [DeliverTx], EndBlock` are renamed to a single method called `FinalizeBlock`.
|
||||
|
||||
#### Summary
|
||||
|
||||
We include a more detailed list of features / scaling improvements that are blocked, and which new phases resolve them at the end of this document.
|
||||
|
||||
<image src="images/abci.png" style="float: left; width: 40%;" /> <image src="images/abci++.png" style="float: right; width: 40%;" />
|
||||
On the top is the existing definition of ABCI, and on the bottom is the proposed ABCI++.
|
||||
|
||||
## Proposal
|
||||
|
||||
Below we suggest an API to add these three new phases.
|
||||
In this document, sometimes the final round of voting is referred to as precommit for clarity in how it acts in the Tendermint case.
|
||||
|
||||
### Prepare Proposal
|
||||
|
||||
*Note, APIs in this section will change after Vote Extensions, we list the adjusted APIs further in the proposal.*
|
||||
|
||||
The Prepare Proposal phase allows the block proposer to perform application-dependent work in a block, to lower the amount of work the rest of the network must do. This enables batch optimizations to a block, which has been empirically demonstrated to be a key component for scaling. This phase introduces the following ABCI method
|
||||
|
||||
```rust
|
||||
fn PrepareProposal(Block) -> BlockData
|
||||
```
|
||||
|
||||
where `BlockData` is a type alias for however data is internally stored within the consensus engine. In Tendermint Core today, this is `[]Tx`.
|
||||
|
||||
The application may read the entire block proposal, and mutate the block data fields. Mutated transactions will still get removed from the mempool later on, as the mempool rechecks all transactions after a block is executed.
|
||||
|
||||
The `PrepareProposal` API will be modified in the vote extensions section, for allowing the application to modify the header.
|
||||
|
||||
### Process Proposal
|
||||
|
||||
The Process Proposal phase sends the block data to the state machine, prior to running the last round of votes on the state machine. This enables features such as allowing validators to reject a block according to whether state machine deems it valid, and changing block execution pipeline.
|
||||
|
||||
We introduce three new methods,
|
||||
|
||||
```rust
|
||||
fn VerifyHeader(header: Header, isValidator: bool) -> ResponseVerifyHeader {...}
|
||||
fn ProcessProposal(block: Block) -> ResponseProcessProposal {...}
|
||||
fn RevertProposal(height: usize, round: usize) {...}
|
||||
```
|
||||
|
||||
where
|
||||
|
||||
```rust
|
||||
struct ResponseVerifyHeader {
|
||||
accept_header: bool,
|
||||
evidence: Vec<Evidence>
|
||||
}
|
||||
struct ResponseProcessProposal {
|
||||
accept_block: bool,
|
||||
evidence: Vec<Evidence>
|
||||
}
|
||||
```
|
||||
|
||||
Upon receiving a block header, every validator runs `VerifyHeader(header, isValidator)`. The reason for why `VerifyHeader` is split from `ProcessProposal` is due to the later sections for Preprocess Proposal and Vote Extensions, where there may be application dependent data in the header that must be verified before accepting the header.
|
||||
If the returned `ResponseVerifyHeader.accept_header` is false, then the validator must precommit nil on this block, and reject all other precommits on this block. `ResponseVerifyHeader.evidence` is appended to the validators local `EvidencePool`.
|
||||
|
||||
Upon receiving an entire block proposal (in the current implementation, all "block parts"), every validator runs `ProcessProposal(block)`. If the returned `ResponseProcessProposal.accept_block` is false, then the validator must precommit nil on this block, and reject all other precommits on this block. `ResponseProcessProposal.evidence` is appended to the validators local `EvidencePool`.
|
||||
|
||||
Once a validator knows that consensus has failed to be achieved for a given block, it must run `RevertProposal(block.height, block.round)`, in order to signal to the application to revert any potentially mutative state changes it may have made. In Tendermint, this occurs when incrementing rounds.
|
||||
|
||||
**RFC**: How do we handle the scenario where honest node A finalized on round x, and honest node B finalized on round x + 1? (e.g. when 2f precommits are publicly known, and a validator precommits themself but doesn't broadcast, but they increment rounds) Is this a real concern? The state root derived could change if everyone finalizes on round x+1, not round x, as the state machine can depend non-uniformly on timestamp.
|
||||
|
||||
The application is expected to cache the block data for later execution.
|
||||
|
||||
The `isValidator` flag is set according to whether the current node is a validator or a full node. This is intended to allow for beginning validator-dependent computation that will be included later in vote extensions. (An example of this is threshold decryptions of ciphertexts.)
|
||||
|
||||
### DeliverTx rename to FinalizeBlock
|
||||
|
||||
After implementing `ProcessProposal`, txs no longer need to be delivered during the block execution phase. Instead, they are already in the state machine. Thus `BeginBlock, DeliverTx, EndBlock` can all be replaced with a single ABCI method for `ExecuteBlock`. Internally the application may still structure its method for executing the block as `BeginBlock, DeliverTx, EndBlock`. However, it is overly restrictive to enforce that the block be executed after it is finalized. There are multiple other, very reasonable pipelined execution models one can go for. So instead we suggest calling this succession of methods `FinalizeBlock`. We propose the following API
|
||||
|
||||
Replace the `BeginBlock, DeliverTx, EndBlock` ABCI methods with the following method
|
||||
|
||||
```rust
|
||||
fn FinalizeBlock() -> ResponseFinalizeBlock
|
||||
```
|
||||
|
||||
where `ResponseFinalizeBlock` has the following API, in terms of what already exists
|
||||
|
||||
```rust
|
||||
struct ResponseFinalizeBlock {
|
||||
updates: ResponseEndBlock,
|
||||
tx_results: Vec<ResponseDeliverTx>
|
||||
}
|
||||
```
|
||||
|
||||
`ResponseEndBlock` should then be renamed to `ConsensusUpdates` and `ResponseDeliverTx` should be renamed to `ResponseTx`.
|
||||
|
||||
### Vote Extensions
|
||||
|
||||
The Vote Extensions phase allow applications to force their validators to do more than just validate within consensus. This is done by allowing the application to add more data to their votes, in the final round of voting. (Namely the precommit)
|
||||
This additional application data will then appear in the block header.
|
||||
|
||||
First we discuss the API changes to the vote struct directly
|
||||
|
||||
```rust
|
||||
fn ExtendVote(height: u64, round: u64) -> (UnsignedAppVoteData, SelfAuthenticatingAppData)
|
||||
fn VerifyVoteExtension(signed_app_vote_data: Vec<u8>, self_authenticating_app_vote_data: Vec<u8>) -> bool
|
||||
```
|
||||
|
||||
There are two types of data that the application can enforce validators to include with their vote.
|
||||
There is data that the app needs the validator to sign over in their vote, and there can be self-authenticating vote data. Self-authenticating here means that the application upon seeing these bytes, knows its valid, came from the validator and is non-malleable. We give an example of each type of vote data here, to make their roles clearer.
|
||||
|
||||
- Unsigned app vote data: A use case of this is if you wanted validator backed oracles, where each validator independently signs some oracle data in their vote, and the median of these values is used on chain. Thus we leverage consensus' signing process for convenience, and use that same key to sign the oracle data.
|
||||
- Self-authenticating vote data: A use case of this is in threshold random beacons. Every validator produces a threshold beacon share. This threshold beacon share can be verified by any node in the network, given the share and the validators public key (which is not the same as its consensus public key). However, this decryption share will not make it into the subsequent block's header. They will be aggregated by the subsequent block proposer to get a single random beacon value that will appear in the subsequent block's header. Everyone can then verify that this aggregated value came from the requisite threshold of the validator set, without increasing the bandwidth for full nodes or light clients. To achieve this goal, the self-authenticating vote data cannot be signed over by the consensus key along with the rest of the vote, as that would require all full nodes & light clients to know this data in order to verify the vote.
|
||||
|
||||
The `CanonicalVote` struct will acommodate the `UnsignedAppVoteData` field by adding another string to its encoding, after the `chain-id`. This should not interfere with existing hardware signing integrations, as it does not affect the constant offset for the `height` and `round`, and the vote size does not have an explicit upper bound. (So adding this unsigned app vote data field is equivalent from the HSM's perspective as having a superlong chain-ID)
|
||||
|
||||
**RFC**: Please comment if you think it will be fine to have elongate the message the HSM signs, or if we need to explore pre-hashing the app vote data.
|
||||
|
||||
The flow of these methods is that when a validator has to precommit, Tendermint will first produce a precommit canonical vote without the application vote data. It will then pass it to the application, which will return unsigned application vote data, and self authenticating application vote data. It will bundle the `unsigned_application_vote_data` into the canonical vote, and pass it to the HSM to sign. Finally it will package the self-authenticating app vote data, and the `signed_vote_data` together, into one final Vote struct to be passed around the network.
|
||||
|
||||
#### Changes to Prepare Proposal Phase
|
||||
|
||||
There are many use cases where the additional data from vote extensions can be batch optimized.
|
||||
This is mainly of interest when the votes include self-authenticating app vote data that be batched together, or the unsigned app vote data is the same across all votes.
|
||||
To allow for this, we change the PrepareProposal API to the following
|
||||
|
||||
```rust
|
||||
fn PrepareProposal(Block, UnbatchedHeader) -> (BlockData, Header)
|
||||
```
|
||||
|
||||
where `UnbatchedHeader` essentially contains a "RawCommit", the `Header` contains a batch-optimized `commit` and an additional "Application Data" field in its root. This will involve a number of changes to core data structures, which will be gone over in the ADR.
|
||||
The `Unbatched` header and `rawcommit` will never be broadcasted, they will be completely internal to consensus.
|
||||
|
||||
#### Inter-process communication (IPC) effects
|
||||
|
||||
For brevity in exposition above, we did not discuss the trade-offs that may occur in interprocess communication delays that these changs will introduce.
|
||||
These new ABCI methods add more locations where the application must communicate with the consensus engine.
|
||||
In most configurations, we expect that the consensus engine and the application will be either statically or dynamically linked, so all communication is a matter of at most adjusting the memory model the data is layed out within.
|
||||
This memory model conversion is typically considered negligible, as delay here is measured on the order of microseconds at most, whereas we face milisecond delays due to cryptography and network overheads.
|
||||
Thus we ignore the overhead in the case of linked libraries.
|
||||
|
||||
In the case where the consensus engine and the application are ran in separate processes, and thus communicate with a form of Inter-process communication (IPC), the delays can easily become on the order of miliseconds based upon the data sent. Thus its important to consider whats happening here.
|
||||
We go through this phase by phase.
|
||||
|
||||
##### Prepare proposal IPC overhead
|
||||
|
||||
This requires a round of IPC communication, where both directions are quite large. Namely the proposer communicating an entire block to the application.
|
||||
However, this can be mitigated by splitting up `PrepareProposal` into two distinct, async methods, one for the block IPC communication, and one for the Header IPC communication.
|
||||
|
||||
Then for chains where the block data does not depend on the header data, the block data IPC communication can proceed in parallel to the prior block's voting phase. (As a node can know whether or not its the leader in the next round)
|
||||
|
||||
Furthermore, this IPC communication is expected to be quite low relative to the amount of p2p gossip time it takes to send the block data around the network, so this is perhaps a premature concern until more sophisticated block gossip protocols are implemented.
|
||||
|
||||
##### Process Proposal IPC overhead
|
||||
|
||||
This phase changes the amount of time available for the consensus engine to deliver a block's data to the state machine.
|
||||
Before, the block data for block N would be delivered to the state machine upon receiving a commit for block N and then be executed.
|
||||
The state machine would respond after executing the txs and before prevoting.
|
||||
The time for block delivery from the consensus engine to the state machine after this change is the time of receiving block proposal N to the to time precommit on proposal N.
|
||||
It is expected that this difference is unimportant in practice, as this time is in parallel to one round of p2p communication for prevoting, which is expected to be significantly less than the time for the consensus engine to deliver a block to the state machine.
|
||||
|
||||
##### Vote Extension IPC overhead
|
||||
|
||||
This has a small amount of data, but does incur an IPC round trip delay. This IPC round trip delay is pretty negligible as compared the variance in vote gossip time. (the IPC delay is typically on the order of 10 microseconds)
|
||||
|
||||
## Status
|
||||
|
||||
Proposed
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- Enables a large number of new features for applications
|
||||
- Supports both immediate and delayed execution models
|
||||
- Allows application specific data from each validator
|
||||
- Allows for batch optimizations across txs, and votes
|
||||
|
||||
### Negative
|
||||
|
||||
- This is a breaking change to all existing ABCI clients, however the application should be able to have a thin wrapper to replicate existing ABCI behavior.
|
||||
- PrepareProposal - can be a no-op
|
||||
- Process Proposal - has to cache the block, but can otherwise be a no-op
|
||||
- Vote Extensions - can be a no-op
|
||||
- Finalize Block - Can black-box call BeginBlock, DeliverTx, EndBlock given the cached block data
|
||||
|
||||
- Vote Extensions adds more complexity to core Tendermint Data Structures
|
||||
- Allowing alternate alternate execution models will lead to a proliferation of new ways for applications to violate expected guarantees.
|
||||
|
||||
### Neutral
|
||||
|
||||
- IPC overhead considerations change, but mostly for the better
|
||||
|
||||
## References
|
||||
|
||||
Reference for IPC delay constants: <http://pages.cs.wisc.edu/~adityav/Evaluation_of_Inter_Process_Communication_Mechanisms.pdf>
|
||||
|
||||
### Short list of blocked features / scaling improvements with required ABCI++ Phases
|
||||
|
||||
| Feature | PrepareProposal | ProcessProposal | Vote Extensions |
|
||||
| :--- | :---: | :---: | :---: |
|
||||
| Tx based signature aggregation | X | | |
|
||||
| SNARK proof of valid state transition | X | | |
|
||||
| Validator provided authentication paths in stateless blockchains | X | | |
|
||||
| Immediate Execution | | X | |
|
||||
| Simple soft forks | | X | |
|
||||
| Validator guaranteed IBC connection attempts | | | X |
|
||||
| Validator based price oracles | | | X |
|
||||
| Immediate Execution with increased time for block execution | X | X | X |
|
||||
| Threshold Encrypted txs | X | X | X |
|
||||
@@ -0,0 +1,94 @@
|
||||
# RFC 014: Semantic Versioning
|
||||
|
||||
## Changelog
|
||||
|
||||
- 2021-11-19: Initial Draft
|
||||
- 2021-02-11: Migrate RFC to tendermint repo (Originally [RFC 006](https://github.com/tendermint/spec/pull/365))
|
||||
|
||||
## Author(s)
|
||||
|
||||
- Callum Waters @cmwaters
|
||||
|
||||
## Context
|
||||
|
||||
We use versioning as an instrument to hold a set of promises to users and signal when such a set changes and how. In the conventional sense of a Go library, major versions signal that the public Go API’s have changed in a breaking way and thus require the users of such libraries to change their usage accordingly. Tendermint is a bit different in that there are multiple users: application developers (both in-process and out-of-process), node operators, and external clients. More importantly, both how these users interact with Tendermint and what's important to these users differs from how users interact and what they find important in a more conventional library.
|
||||
|
||||
This document attempts to encapsulate the discussions around versioning in Tendermint and draws upon them to propose a guide to how Tendermint uses versioning to make promises to its users.
|
||||
|
||||
For a versioning policy to make sense, we must also address the intended frequency of breaking changes. The strictest guarantees in the world will not help users if we plan to break them with every release.
|
||||
|
||||
Finally I would like to remark that this RFC only addresses the "what", as in what are the rules for versioning. The "how" of Tendermint implementing the versioning rules we choose, will be addressed in a later RFC on Soft Upgrades.
|
||||
|
||||
## Discussion
|
||||
|
||||
We first begin with a round up of the various users and a set of assumptions on what these users expect from Tendermint in regards to versioning:
|
||||
|
||||
1. **Application Developers**, those that use the ABCI to build applications on top of Tendermint, are chiefly concerned with that API. Breaking changes will force developers to modify large portions of their codebase to accommodate for the changes. Some ABCI changes such as introducing priority for the mempool don't require any effort and can be lazily adopted whilst changes like ABCI++ may force applications to redesign their entire execution system. It's also worth considering that the API's for go developers differ to developers of other languages. The former here can use the entire Tendermint library, most notably the local RPC methods, and so the team must be wary of all public Go API's.
|
||||
2. **Node Operators**, those running node infrastructure, are predominantly concerned with downtime, complexity and frequency of upgrading, and avoiding data loss. They may be also concerned about changes that may break the scripts and tooling they use to supervise their nodes.
|
||||
3. **External Clients** are those that perform any of the following:
|
||||
- consume the RPC endpoints of nodes like `/block`
|
||||
- subscribe to the event stream
|
||||
- make queries to the indexer
|
||||
|
||||
This set are concerned with chain upgrades which will impact their ability to query state and block data as well as broadcast transactions. Examples include wallets and block explorers.
|
||||
|
||||
4. **IBC module and relayers**. The developers of IBC and consumers of their software are concerned about changes that may affect a chain's ability to send arbitrary messages to another chain. Specifically, these users are affected by any breaking changes to the light client verification algorithm.
|
||||
|
||||
Although we present them here as having different concerns, in a broader sense these user groups share a concern for the end users of applications. A crucial principle guiding this RFC is that **the ability for chains to provide continual service is more important than the actual upgrade burden put on the developers of these chains**. This means some extra burden for application developers is tolerable if it minimizes or substantially reduces downtime for the end user.
|
||||
|
||||
### Modes of Interprocess Communication
|
||||
|
||||
Tendermint has two primary mechanisms to communicate with other processes: RPC and P2P. The division marks the boundary between the internal and external components of the network:
|
||||
|
||||
- The P2P layer is used in all cases that nodes (of any type) need to communicate with one another.
|
||||
- The RPC interface is for any outside process that wants to communicate with a node.
|
||||
|
||||
The design principle here is that **communication via RPC is to a trusted source** and thus the RPC service prioritizes inspection rather than verification. The P2P interface is the primary medium for verification.
|
||||
|
||||
As an example, an in-browser light client would verify headers (and perhaps application state) via the p2p layer, and then pass along information on to the client via RPC (or potentially directly via a separate API).
|
||||
|
||||
The main exceptions to this are the IBC module and relayers, which are external to the node but also require verifiable data. Breaking changes to the light client verification path mean that all neighbouring chains that are connected will no longer be able to verify state transitions and thus pass messages back and forward.
|
||||
|
||||
## Proposal
|
||||
|
||||
Tendermint version labels will follow the syntax of [Semantic Versions 2.0.0](https://semver.org/) with a major, minor and patch version. The version components will be interpreted according to these rules:
|
||||
|
||||
For the entire cycle of a **major version** in Tendermint:
|
||||
|
||||
- All blocks and state data in a blockchain can be queried. All headers can be verified even across minor version changes. Nodes can both block sync and state sync from genesis to the head of the chain.
|
||||
- Nodes in a network are able to communicate and perform BFT state machine replication so long as the agreed network version is the lowest of all nodes in a network. For example, nodes using version 1.5.x and 1.2.x can operate together so long as the network version is 1.2 or lower (but still within the 1.x range). This rule essentially captures the concept of network backwards compatibility.
|
||||
- Node RPC endpoints will remain compatible with existing external clients:
|
||||
- New endpoints may be added, but old endpoints may not be removed.
|
||||
- Old endpoints may be extended to add new request and response fields, but requests not using those fields must function as before the change.
|
||||
- Migrations should be automatic. Upgrading of one node can happen asynchronously with respect to other nodes (although agreement of a network-wide upgrade must still occur synchronously via consensus).
|
||||
|
||||
For the entire cycle of a **minor version** in Tendermint:
|
||||
|
||||
- Public Go API's, for example in `node` or `abci` packages will not change in a way that requires any consumer (not just application developers) to modify their code.
|
||||
- No breaking changes to the block protocol. This means that all block related data structures should not change in a way that breaks any of the hashes, the consensus engine or light client verification.
|
||||
- Upgrades between minor versions may not result in any downtime (i.e., no migrations are required), nor require any changes to the config files to continue with the existing behavior. A minor version upgrade will require only stopping the existing process, swapping the binary, and starting the new process.
|
||||
|
||||
A new **patch version** of Tendermint will only contain bug fixes and updates that impact the security and stability of Tendermint.
|
||||
|
||||
These guarantees will come into effect at release 1.0.
|
||||
|
||||
## Status
|
||||
|
||||
Proposed
|
||||
|
||||
## Consequences
|
||||
|
||||
### Positive
|
||||
|
||||
- Clearer communication of what versioning means to us and the effect they have on our users.
|
||||
|
||||
### Negative
|
||||
|
||||
- Can potentially incur greater engineering effort to uphold and follow these guarantees.
|
||||
|
||||
### Neutral
|
||||
|
||||
## References
|
||||
|
||||
- [SemVer](https://semver.org/)
|
||||
- [Tendermint Tracking Issue](https://github.com/tendermint/tendermint/issues/5680)
|
||||
@@ -0,0 +1,261 @@
|
||||
# RFC 015: ABCI++ TX Mutation
|
||||
|
||||
## Changelog
|
||||
|
||||
- 23-Feb-2022: Initial draft (@williambanfield).
|
||||
- 28-Feb-2022: Revised draft (@williambanfield).
|
||||
|
||||
## Abstract
|
||||
|
||||
A previous version of the ABCI++ specification detailed a mechanism for proposers to replace transactions
|
||||
in the proposed block. This scheme required the proposer to construct new transactions
|
||||
and mark these new transactions as replacing other removed transactions. The specification
|
||||
was ambiguous as to how the replacement may be communicated to peer nodes.
|
||||
This RFC discusses issues with this mechanism and possible solutions.
|
||||
|
||||
## Background
|
||||
|
||||
### What is the proposed change?
|
||||
|
||||
A previous version of the ABCI++ specification proposed mechanisms for adding, removing, and replacing
|
||||
transactions in a proposed block. To replace a transaction, the application running
|
||||
`ProcessProposal` could mark a transaction as replaced by other application-supplied
|
||||
transactions by returning a new transaction marked with the `ADDED` flag setting
|
||||
the `new_hashes` field of the removed transaction to contain the list of transaction hashes
|
||||
that replace it. In that previous specification for ABCI++, the full use of the
|
||||
`new_hashes` field is left somewhat ambiguous. At present, these hashes are not
|
||||
gossiped and are not eventually included in the block to signal replacement to
|
||||
other nodes. The specification did indicate that the transactions specified in
|
||||
the `new_hashes` field will be removed from the mempool but it's not clear how
|
||||
peer nodes will learn about them.
|
||||
|
||||
### What systems would be affected by adding transaction replacement?
|
||||
|
||||
The 'transaction' is a central building block of a Tendermint blockchain, so adding
|
||||
a mechanism for transaction replacement would require changes to many aspects of Tendermint.
|
||||
|
||||
The following is a rough list of the functionality that this mechanism would affect:
|
||||
|
||||
#### Transaction indexing
|
||||
|
||||
Tendermint's indexer stores transactions and transaction results using the hash of the executed
|
||||
transaction [as the key][tx-result-index] and the ABCI results and transaction bytes as the value.
|
||||
|
||||
To allow transaction replacement, the replaced transactions would need to stored as well in the
|
||||
indexer, likely as a mapping of original transaction to list of transaction hashes that replaced
|
||||
the original transaction.
|
||||
|
||||
#### Transaction inclusion proofs
|
||||
|
||||
The result of a transaction query includes a Merkle proof of the existence of the
|
||||
transaction in the block chain. This [proof is built][inclusion-proof] as a merkle tree
|
||||
of the hashes of all of the transactions in the block where the queried transaction was executed.
|
||||
|
||||
To allow transaction replacement, these proofs would need to be updated to prove
|
||||
that a replaced transaction was included by replacement in the block.
|
||||
|
||||
#### RPC-based transaction query parameters and results
|
||||
|
||||
Tendermint's RPC allows clients to retrieve information about transactions via the
|
||||
`/tx_search` and `/tx` RPC endpoints.
|
||||
|
||||
RPC query results containing replaced transactions would need to be updated to include
|
||||
information on replaced transactions, either by returning results for all of the replaced
|
||||
transactions, or by including a response with just the hashes of the replaced transactions
|
||||
which clients could proceed to query individually.
|
||||
|
||||
#### Mempool transaction removal
|
||||
|
||||
Additional logic would need to be added to the Tendermint mempool to clear out replaced
|
||||
transactions after each block is executed. Tendermint currently removes executed transactions
|
||||
from the mempool, so this would be a pretty straightforward change.
|
||||
|
||||
## Discussion
|
||||
|
||||
### What value may be added to Tendermint by introducing transaction replacement?
|
||||
|
||||
Transaction replacement would would enable applications to aggregate or disaggregate transactions.
|
||||
|
||||
For aggregation, a set of transactions that all related work, such as transferring
|
||||
tokens between the same two accounts, could be replaced with a single transaction,
|
||||
i.e. one that transfers a single sum from one account to the other.
|
||||
Applications that make frequent use of aggregation may be able to achieve a higher throughput.
|
||||
Aggregation would decrease the space occupied by a single client-submitted transaction in the block, allowing
|
||||
more client-submitted transactions to be executed per block.
|
||||
|
||||
For disaggregation, a very complex transaction could be split into multiple smaller transactions.
|
||||
This may be useful if an application wishes to perform more fine-grained indexing on intermediate parts
|
||||
of a multi-part transaction.
|
||||
|
||||
### Drawbacks to transaction replacement
|
||||
|
||||
Transaction replacement would require updating and shimming many of the places that
|
||||
Tendermint records and exposes information about executed transactions. While
|
||||
systems within Tendermint could be updated to account for transaction replacement,
|
||||
such a system would leave new issues and rough edges.
|
||||
|
||||
#### No way of guaranteeing correct replacement
|
||||
|
||||
If a user issues a transaction to the network and the transaction is replaced, the
|
||||
user has no guarantee that the replacement was correct. For example, suppose a set of users issue
|
||||
transactions A, B, and C and they are all aggregated into a new transaction, D.
|
||||
There is nothing guaranteeing that D was constructed correctly from the inputs.
|
||||
The only way for users to ensure D is correct would be if D contained all of the
|
||||
information of its constituent transactions, in which case, nothing is really gained by the replacement.
|
||||
|
||||
#### Replacement transactions not signed by submitter
|
||||
|
||||
Abstractly, Tendermint simply views transactions as a ball of bytes and therefore
|
||||
should be fine with replacing one for another. However, many applications require
|
||||
that transactions submitted to the chain be signed by some private key to authenticate
|
||||
and authorize the transaction. Replaced transactions could not be signed by the
|
||||
submitter, only by the application node. Therefore, any use of transaction replacement
|
||||
could not contain authorization from the submitter and would either need to grant
|
||||
application-submitted transactions power to perform application logic on behalf
|
||||
of a user without their consent.
|
||||
|
||||
Granting this power to application-submitted transactions would be very dangerous
|
||||
and therefore might not be of much value to application developers.
|
||||
Transaction replacement might only be really safe in the case of application-submitted
|
||||
transactions or for transactions that require no authorization. For such transactions,
|
||||
it's quite not quite clear what the utility of replacement is: the application can already
|
||||
generate any transactions that it wants. The fact that such a transaction was a replacement
|
||||
is not particularly relevant to participants in the chain since the application is
|
||||
merely replacing its own transactions.
|
||||
|
||||
#### New vector for censorship
|
||||
|
||||
Depending on the implementation, transaction replacement may allow a node signal
|
||||
to the rest of the chain that some transaction should no longer be considered for execution.
|
||||
Honest nodes will use the replacement mechanism to signal that a transaction has been aggregated.
|
||||
Malicious nodes will be granted a new vector for censoring transactions.
|
||||
There is no guarantee that a replaced transactions is actually executed at all.
|
||||
A malicious node could censor a transaction by simply listing it as replaced.
|
||||
Honest nodes seeing the replacement would flush the transaction from their mempool
|
||||
and not execute or propose it it in later blocks.
|
||||
|
||||
### Transaction tracking implementations
|
||||
|
||||
This section discusses possible ways to flesh out the implementation of transaction replacement.
|
||||
Specifically, this section proposes a few alternative ways that Tendermint blockchains could
|
||||
track and store transaction replacements.
|
||||
|
||||
#### Include transaction replacements in the block
|
||||
|
||||
One option to track transaction replacement is to include information on the
|
||||
transaction replacement within the block. An additional structure may be added
|
||||
the block of the following form:
|
||||
|
||||
```proto
|
||||
message Block {
|
||||
...
|
||||
repeated Replacement replacements = 5;
|
||||
}
|
||||
|
||||
message Replacement {
|
||||
bytes included_tx_key = 1;
|
||||
repeated bytes replaced_txs_keys = 2;
|
||||
}
|
||||
```
|
||||
|
||||
Applications executing `PrepareProposal` would return the list of replacements and
|
||||
Tendermint would include an encoding of these replacements in the block that is gossiped
|
||||
and committed.
|
||||
|
||||
Tendermint's transaction indexing would include a new mapping for each replaced transaction
|
||||
key to the committed transaction.
|
||||
Transaction inclusion proofs would be updated to include these additional new transaction
|
||||
keys in the Merkle tree and queries for transaction hashes that were replaced would return
|
||||
information indicating that the transaction was replaced along with the hash of the
|
||||
transaction that replaced it.
|
||||
|
||||
Block validation of gossiped blocks would be updated to check that each of the
|
||||
`included_txs_key` matches the hash of some transaction in the proposed block.
|
||||
|
||||
Implementing the changes described in this section would allow Tendermint to gossip
|
||||
and index transaction replacements as part of block propagation. These changes would
|
||||
still require the application to certify that the replacements were valid. This
|
||||
validation may be performed in one of two ways:
|
||||
|
||||
1. **Applications optimistically trust that the proposer performed a legitimate replacement.**
|
||||
|
||||
In this validation scheme, applications would not verify that the substitution
|
||||
is valid during consensus and instead simply trust that the proposer is correct.
|
||||
This would have the drawback of allowing a malicious proposer to remove transactions
|
||||
it did not want executed.
|
||||
|
||||
2. **Applications completely validate transaction replacement.**
|
||||
|
||||
In this validation scheme, applications that allow replacement would check that
|
||||
each listed replaced transaction was correctly reflected in the replacement transaction.
|
||||
In order to perform such validation, the node would need to have the replaced transactions
|
||||
locally. This could be accomplished one of a few ways: by querying the mempool,
|
||||
by adding an additional p2p gossip channel for transaction replacements, or by including the replaced transactions
|
||||
in the block. Replacement validation via mempool querying would require the node
|
||||
to have received all of the replaced transactions in the mempool which is far from
|
||||
guaranteed. Adding an additional gossip channel would make gossiping replaced transactions
|
||||
a requirement for consensus to proceed, since all nodes would need to receive all replacement
|
||||
messages before considering a block valid. Finally, including replaced transactions in
|
||||
the block seems to obviate any benefit gained from performing a transaction replacement
|
||||
since the replaced transaction and the original transactions would now both appear in the block.
|
||||
|
||||
#### Application defined transaction replacement
|
||||
|
||||
An additional option for allowing transaction replacement is to leave it entirely as a responsibility
|
||||
of the application. The `PrepareProposal` ABCI++ call allows for applications to add
|
||||
new transactions to a proposed block. Applications that wished to implement a transaction
|
||||
replacement mechanism would be free to do so without the newly defined `new_hashes` field.
|
||||
Applications wishing to implement transaction replacement would add the aggregated
|
||||
transactions in the `PrepareProposal` response, and include one additional bookkeeping
|
||||
transaction that listed all of the replacements, with a similar scheme to the `new_hashes`
|
||||
field described in ABCI++. This new bookkeeping transaction could be used by the
|
||||
application to determine which transactions to clear from the mempool in future calls
|
||||
to `CheckTx`.
|
||||
|
||||
The meaning of any transaction in the block is completely opaque to Tendermint,
|
||||
so applications performing this style of replacement would not be able to have the replacement
|
||||
reflected in any most of Tendermint's transaction tracking mechanisms, such as transaction indexing
|
||||
and the `/tx` endpoint.
|
||||
|
||||
#### Application defined Tx Keys
|
||||
|
||||
Tendermint currently uses cryptographic hashes, SHA256, as a key for each transaction.
|
||||
As noted in the section on systems that would require changing, this key is used
|
||||
to identify the transaction in the mempool, in the indexer, and within the RPC system.
|
||||
|
||||
An alternative approach to allowing `ProcessProposal` to specify a set of transaction
|
||||
replacements would be instead to allow the application to specify an additional key or set
|
||||
of keys for each transaction during `ProcessProposal`. This new `secondary_keys` set
|
||||
would be included in the block and therefore gossiped during block propagation.
|
||||
Additional RPC endpoints could be exposed to query by the application-defined keys.
|
||||
|
||||
Applications wishing to implement replacement would leverage this new field by providing the
|
||||
replaced transaction hashes as the `secondary_keys` and checking their validity during
|
||||
`ProcessProposal`. During `RecheckTx` the application would then be responsible for
|
||||
clearing out transactions that matched the `secondary_keys`.
|
||||
|
||||
It is worth noting that something like this would be possible without `secondary_keys`.
|
||||
An application wishing to implement a system like this one could define a replacement
|
||||
transaction, as discussed in the section on application-defined transaction replacement,
|
||||
and use a custom [ABCI event type][abci-event-type] to communicate that the replacement should
|
||||
be indexed within Tendermint's ABCI event indexing.
|
||||
|
||||
### Complexity to value-add tradeoff
|
||||
|
||||
It is worth remarking that adding a system like this may introduce a decent amount
|
||||
of new complexity into Tendermint. An approach that leaves much of the replacement
|
||||
logic to Tendermint would require altering the core transaction indexing and querying
|
||||
data. In many of the cases listed, a system for transaction replacement is possible
|
||||
without explicitly defining it as part of `PrepareProposal`. Since applications
|
||||
can now add transactions during `PrepareProposal` they can and should leverage this
|
||||
functionality to include additional bookkeeping transactions in the block. It may
|
||||
be worth encouraging applications to discover new and interesting ways to leverage this
|
||||
power instead of immediately solving the problem for them.
|
||||
|
||||
### References
|
||||
|
||||
[inclusion-proof]: https://github.com/tendermint/tendermint/blob/0fcfaa4568cb700e27c954389c1fcd0b9e786332/types/tx.go#L67
|
||||
[tx-serach-result]: https://github.com/tendermint/tendermint/blob/0fcfaa4568cb700e27c954389c1fcd0b9e786332/rpc/coretypes/responses.go#L267
|
||||
[tx-rpc-func]: https://github.com/tendermint/tendermint/blob/0fcfaa4568cb700e27c954389c1fcd0b9e786332/internal/rpc/core/tx.go#L21
|
||||
[tx-result-index]: https://github.com/tendermint/tendermint/blob/0fcfaa4568cb700e27c954389c1fcd0b9e786332/internal/state/indexer/tx/kv/kv.go#L90
|
||||
[abci-event-type]: https://github.com/tendermint/tendermint/blob/0fcfaa4568cb700e27c954389c1fcd0b9e786332/abci/types/types.pb.go#L3168
|
||||
@@ -0,0 +1,83 @@
|
||||
# RFC 016: Node Architecture
|
||||
|
||||
## Changelog
|
||||
|
||||
- April 8, 2022: Initial draft (@cmwaters)
|
||||
- April 15, 2022: Incorporation of feedback
|
||||
|
||||
## Abstract
|
||||
|
||||
The `node` package is the entry point into the Tendermint codebase, used both by the command line and programatically to create the nodes that make up a network. The package has suffered the most from the evolution of the codebase, becoming bloated as developers clipped on their bits of code here and there to get whatever feature they wanted working.
|
||||
|
||||
The decisions made at the node level have the biggest impact to simplifying the protocols within them, unlocking better internal designs and making Tendermint more intuitive to use and easier to understand from the outside. Work, in minor increments, has already begun on this section of the codebase. This document exists to spark forth the necessary discourse in a few related areas that will help the team to converge on the long term makeup of the node.
|
||||
|
||||
## Discussion
|
||||
|
||||
The following is a list of points of discussion around the architecture of the node:
|
||||
|
||||
### Dependency Tree
|
||||
|
||||
The node object is currently stuffed with every component that possibly exists within Tendermint. In the constructor, all objects are built and interlaid with one another in some awkward dance. My guiding principle is that the node should only be made up of the components that it wants to have direct control of throughout its life. The node is a service which currently has the purpose of starting other services up in a particular order and stopping them all when commanded to do so. However, there are many services which are not direct dependents i.e. the mempool and evidence services should only be working when the consensus service is running. I propose to form more of a hierarchical structure of dependents which forces us to be clear about the relations that one component has to the other. More concretely, I propose the following dependency tree:
|
||||
|
||||

|
||||
|
||||
Many of the further discussion topics circle back to this representation of the node.
|
||||
|
||||
It's also important to distinguish two dimensions which may require different characteristics of the architecture. There is the starting and stopping of services and their general lifecycle management. What is the correct order of operations to starting a node for example. Then there is the question of the needs of the service during actual operation. Then there is the question of what resources each service needs access to during its operation. Some need to publish events, others need access to data stores, and so forth.
|
||||
|
||||
An alternative model and one that perhaps better suits the latter of these dimensions is the notion of an internal message passing system. Either the events bus or p2p layer could serve as a viable transport. This would essentially allow all services to communicate with any other service and could perhaps provide a solution to the coordination problem (presented below) without a centralized coordinator. The other main advantage is that such a system would be more robust to disruptions and changes to the code which may make a hierarchical structure quickly outdated and suboptimal. The addition of message routing is an added complexity to implement, will increase the degree of asynchronicity in the system and may make it harder to debug problems that are across multiple services.
|
||||
|
||||
### Coordination of State Advancing Mechanisms
|
||||
|
||||
Advancement of state in Tendermint is simply defined in heights: If the node is at height n, how does it get to height n + 1 and so on. Based on this definition we have three components that help a node to advance in height: consensus, statesync and blocksync. The way these components behave currently is very tightly coupled to one another with references passed back and forth. My guiding principle is that each of these should be able to operate completely independently of each other, e.g. a node should be able to run solely blocksync indefinitely. There have been several ideas suggested towards improving this flow. I've been leaning strongly towards a centralized system, whereby an orchestrator (in this case the node) decides what services to start and stop.
|
||||
In a decentralized message passing system, individual services make their decision based upon a "global" shared state i.e. if my height is less that 10 below the average peer height, I as consensus, should stop (knowing that blocksync has the same condition for starting). As the example illustrates, each mechanism will still need to be aware of the presence of other mechanisms.
|
||||
|
||||
Both centralized and decentralized systems rely on the communication of the nodes current height and a judgement on the height of the head of the chain. The latter, working out the head of the chain, is quite a difficult challenge as their is nothing preventing the node from acting maliciously and providing a different height. Currently both blocksync, consensus (and to a certain degree statesync), have parallel systems where peers communicate their height. This could be streamlined with the consensus (or even the p2p layer), broadcasting peer heights and either the node or the other state advancing mechanisms acting accordingly.
|
||||
|
||||
Currently, when a node starts, it turns on every service that it is attached to. This means that while a node is syncing up by requesting blocks, it is also receiving transactions and votes, as well as snapshot and block requests. This is a needless use of bandwidth. An implementation of an orchestrator, regardless of whether the system is heirachical or not, should look to be able to open and close channels dynamically and effectively broadcast which services it is running. Integrating this with service discovery may also lead to a better serivce to peers.
|
||||
|
||||
The orchestrator allows for some deal of variablity in how a node is constructed. Does it just run blocksync, shadowing the head of the chain and be highly available for querying. Does it rely on state sync at all? An important question that arises from this dynamicism is we ideally want to encourage nodes to provide as much of their resources as possible so that their is a healthy amount of providers to consumers. Do we make all services compulsory or allow for them to be disabled? Arguably it's possible that a user forks the codebase and rips out the blocksync code because they want to reduce bandwidth so this is more a question of how easy do we want to make this for users.
|
||||
|
||||
### Block Executor
|
||||
|
||||
The block executor is an important component that is currently used by both consensus and blocksync to execute transactions and update application state. Principally, I think it should be the only component that can write (and possibly even read) the block and state stores, and we should clean up other direct dependencies on the storage engine if we can. This would mean:
|
||||
|
||||
- The reactors Consensus, BlockSync and StateSync should all import the executor for advancing state ie. `ApplyBlock` and `BootstrapState`.
|
||||
- Pruning should also be a concern of the block executor as well as `FinalizeBlock` and `Commit`. This can simplify consensus to focus just on the consensus part.
|
||||
|
||||
### The Interprocess communication systems: RPC, P2P, ABCI, and Events
|
||||
|
||||
The schematic supplied above shows the relations between the different services, the node, the block executor, and the storage layer. Represented as colored dots are the components responsible for different roles of interprocess communication (IPC). These components permeate throughout the code base, seeping into most services. What can provide powerful functionality on one hand can also become a twisted vine, creating messy corner cases and convoluting the protocols themselves. A lot of the thinking around
|
||||
how we want our IPC systens to function has been summarised in this [RFC](./rfc-002-ipc-ecosystem.md). In this section, I'd like to focus the reader on the relation between the IPC and the node structure. An issue that has frequently risen is that the RPC has control of the components where it strikes me as being more logical for the component to dictate the information that is emitted/available and the knobs it wishes to expose. The RPC is also inextricably tied to the node instance and has situations where it is passed pointers directly to the storage engine and other components.
|
||||
|
||||
I am currently convinced of the approach that the p2p layer takes and would like to see other IPC components follow suit. This would mean that the RPC and events system would be constructed in the node yet would pass the adequate methods to register endpoints and topics to the sub components. For example,
|
||||
|
||||
```go
|
||||
// Methods from the RPC and event bus that would be passed into the constructor of components like "consensus"
|
||||
// NOTE: This is a hypothetical construction to convey the idea. An actual implementation may differ.
|
||||
func RegisterRoute(path string, handler func(http.ResponseWriter, *http.Request))
|
||||
|
||||
func RegisterTopic(name string) EventPublisher
|
||||
|
||||
type EventPublisher func (context.Context, types.EventData, []abci.Event)
|
||||
```
|
||||
|
||||
This would give the components control to the information they want to expose and keep all relevant logic within that package. It accomodates more to a dynamic system where services can switch on and off. Each component would also receive access to the logger and metrics system for introspection and debuggability.
|
||||
|
||||
#### IPC Rubric
|
||||
|
||||
I'd like to aim to reach a state where we as a team have either an implicit or explicit rubric which can determine, in the event of some new need to communicate information, what tool it should use for doing this. In the case of inter node communication, this is obviously the p2p stack (with perhaps the exception of the light client). Metrics and logging also have clear usage patterns. RPC and the events system are less clear. The RPC is used for debugging data and fine tuned operator control as it is for general public querying and transaction submission. The RPC is also known to have been plumbed back into the application for historical queries. The events system, similarly, is used for consuming transaction events as it is for the testing of consensus state transitions.
|
||||
|
||||
Principally, I think we should look to change our language away from what the actual transport is and more towards what it's being used for and to whom. We call it a peer to peer layer and not the underlying tcp connection. In the same way, we should look to split RPC into an operator interface (RPC Internal), a public interface (RPC External) and a bidirectional ABCI.
|
||||
|
||||
### Seperation of consumers and suppliers
|
||||
|
||||
When a service such as blocksync is turned on, it automatically begins requesting blocks to verify and apply them as it also tries to serve them to other peers catching up. We should look to distinguish these two aspects: supplying of information and consuming of information in many of these components. More concretely, I'd suggest:
|
||||
|
||||
- The blocksync and statesync service, i.e. supplying information for those trying to catch up should only start running once a node has caught up i.e. after running the blocksync and/or state sync *processes*
|
||||
- The blocksync and state sync processes have defined termination clauses that inform the orchestrator when they are done and where they finished.
|
||||
- One way of achieving this would be that every process both passes and returns the `State` object
|
||||
- In some cases, a node may specify that it wants to run blocksync indefinitely.
|
||||
- The mempool should also indicate whether it wants to receive transactions or to send them only (one-directional mempool)
|
||||
- Similarly, the light client itself only requests information whereas the light client service (currently part of state sync) can do both.
|
||||
- This distinction needs to be communicated in the p2p layer handshake itself but should also be changeable over the lifespan of the connection.
|
||||
@@ -0,0 +1,571 @@
|
||||
# RFC 017: ABCI++ Vote Extension Propagation
|
||||
|
||||
## Changelog
|
||||
|
||||
- 11-Apr-2022: Initial draft (@sergio-mena).
|
||||
- 15-Apr-2022: Addressed initial comments. First complete version (@sergio-mena).
|
||||
- 09-May-2022: Addressed all outstanding comments.
|
||||
|
||||
## Abstract
|
||||
|
||||
According to the
|
||||
[ABCI++ specification](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/abci%2B%2B/README.md)
|
||||
(as of 11-Apr-2022), a validator MUST provide a signed vote extension for each non-`nil` precommit vote
|
||||
of height *h* that it uses to propose a block in height *h+1*. When a validator is up to
|
||||
date, this is easy to do, but when a validator needs to catch up this is far from trivial as this data
|
||||
cannot be retrieved from the blockchain.
|
||||
|
||||
This RFC presents and compares the different options to address this problem, which have been proposed
|
||||
in several discussions by the Tendermint Core team.
|
||||
|
||||
## Document Structure
|
||||
|
||||
The RFC is structured as follows. In the [Background](#background) section,
|
||||
subsections [Problem Description](#problem-description) and [Cases to Address](#cases-to-address)
|
||||
explain the problem at hand from a high level perspective, i.e., abstracting away from the current
|
||||
Tendermint implementation. In contrast, subsection
|
||||
[Current Catch-up Mechanisms](#current-catch-up-mechanisms) delves into the details of the current
|
||||
Tendermint code.
|
||||
|
||||
In the [Discussion](#discussion) section, subsection [Solutions Proposed](#solutions-proposed) is also
|
||||
worded abstracting away from implementation details, whilst subsections
|
||||
[Feasibility of the Proposed Solutions](#feasibility-of-the-proposed-solutions) and
|
||||
[Current Limitations and Possible Implementations](#current-limitations-and-possible-implementations)
|
||||
analize the viability of one of the proposed solutions in the context of Tendermint's architecture
|
||||
based on reactors. Finally, [Formalization Work](#formalization-work) briefly discusses the work
|
||||
still needed demonstrate the correctness of the chosen solution.
|
||||
|
||||
The high level subsections are aimed at readers who are familiar with consensus algorithms, in
|
||||
particular with the one described in the Tendermint (white paper), but who are not necessarily
|
||||
acquainted with the details of the Tendermint codebase. The other subsections, which go into
|
||||
implementation details, are best understood by engineers with deep knowledge of the implementation of
|
||||
Tendermint's blocksync and consensus reactors.
|
||||
|
||||
## Background
|
||||
|
||||
### Basic Definitions
|
||||
|
||||
This document assumes that all validators have equal voting power for the sake of simplicity. This is done
|
||||
without loss of generality.
|
||||
|
||||
There are two types of votes in Tendermint: *prevotes* and *precommits*. Votes can be `nil` or refer to
|
||||
a proposed block. This RFC focuses on precommits,
|
||||
also known as *precommit votes*. In this document we sometimes call them simply *votes*.
|
||||
|
||||
Validators send precommit votes to their peer nodes in *precommit messages*. According to the
|
||||
[ABCI++ specification](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/abci%2B%2B/README.md),
|
||||
a precommit message MUST also contain a *vote extension*.
|
||||
This mandatory vote extension can be empty, but MUST be signed with the same key as the precommit
|
||||
vote (i.e., the sending validator's).
|
||||
Nevertheless, the vote extension is signed independently from the vote, so a vote can be separated from
|
||||
its extension.
|
||||
The reason for vote extensions to be mandatory in precommit messages is that, otherwise, a (malicious)
|
||||
node can omit a vote extension while still providing/forwarding/sending the corresponding precommit vote.
|
||||
|
||||
The validator set at height *h* is denoted *valset<sub>h</sub>*. A *commit* for height *h* consists of more
|
||||
than *2n<sub>h</sub>/3* precommit votes voting for a block *b*, where *n<sub>h</sub>* denotes the size of
|
||||
*valset<sub>h</sub>*. A commit does not contain `nil` precommit votes, and all votes in it refer to the
|
||||
same block. An *extended commit* is a *commit* where every precommit vote has its respective vote extension
|
||||
attached.
|
||||
|
||||
### Problem Description
|
||||
|
||||
In the version of [ABCI](https://github.com/tendermint/spec/blob/4fb99af/spec/abci/README.md) present up to
|
||||
Tendermint v0.35, for any height *h*, a validator *v* MUST have the decided block *b* and a commit for
|
||||
height *h* in order to decide at height *h*. Then, *v* just needs a commit for height *h* to propose at
|
||||
height *h+1*, in the rounds of *h+1* where *v* is a proposer.
|
||||
|
||||
In [ABCI++](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/abci%2B%2B/README.md),
|
||||
the information that a validator *v* MUST have to be able to decide in *h* does not change with
|
||||
respect to pre-existing ABCI: the decided block *b* and a commit for *h*.
|
||||
In contrast, for proposing in *h+1*, a commit for *h* is not enough: *v* MUST now have an extended
|
||||
commit.
|
||||
|
||||
When a validator takes an active part in consensus at height *h*, it has all the data it needs in memory,
|
||||
in its consensus state, to decide on *h* and propose in *h+1*. Things are not so easy in the cases when
|
||||
*v* cannot take part in consensus because it is late (e.g., it falls behind, it crashes
|
||||
and recovers, or it just starts after the others). If *v* does not take part, it cannot actively
|
||||
gather precommit messages (which include vote extensions) in order to decide.
|
||||
Before ABCI++, this was not a problem: full nodes are supposed to persist past blocks in the block store,
|
||||
so other nodes would realise that *v* is late and send it the missing decided block at height *h* and
|
||||
the corresponding commit (kept in block *h+1*) so that *v* can catch up.
|
||||
However, we cannot apply this catch-up technique for ABCI++, as the vote extensions, which are part
|
||||
of the needed *extended commit* are not part of the blockchain.
|
||||
|
||||
### Cases to Address
|
||||
|
||||
Before we tackle the description of the possible cases we need to address, let us describe the following
|
||||
incremental improvement to the ABCI++ logic. Upon decision, a full node persists (e.g., in the block
|
||||
store) the extended commit that allowed the node to decide. For the moment, let us assume the node only
|
||||
needs to keep its *most recent* extended commit, and MAY remove any older extended commits from persistent
|
||||
storage.
|
||||
This improvement is so obvious that all solutions described in the [Discussion](#discussion) section use
|
||||
it as a building block. Moreover, it completely addresses by itself some of the cases described in this
|
||||
subsection.
|
||||
|
||||
We now describe the cases (i.e. possible *runs* of the system) that have been raised in different
|
||||
discussions and need to be addressed. They are (roughly) ordered from easiest to hardest to deal with.
|
||||
|
||||
- **(a)** *Happy path: all validators advance together, no crash*.
|
||||
|
||||
This case is included for completeness. All validators have taken part in height *h*.
|
||||
Even if some of them did not manage to send a precommit message for the decided block, they all
|
||||
receive enough precommit messages to be able to decide. As vote extensions are mandatory in
|
||||
precommit messages, every validator *v* trivially has all the information, namely the decided block
|
||||
and the extended commit, needed to propose in height *h+1* for the rounds in which *v* is the
|
||||
proposer.
|
||||
|
||||
No problem to solve here.
|
||||
|
||||
- **(b)** *All validators advance together, then all crash at the same height*.
|
||||
|
||||
This case has been raised in some discussions, the main concern being whether the vote extensions
|
||||
for the previous height would be lost across the network. With the improvement described above,
|
||||
namely persisting the latest extended commit at decision time, this case is solved.
|
||||
When a crashed validator recovers, it recovers the last extended commit from persistent storage
|
||||
and handshakes with the Application.
|
||||
If need be, it also reconstructs messages for the unfinished height
|
||||
(including all precommits received) from the WAL.
|
||||
Then, the validator can resume where it was at the time of the crash. Thus, as extensions are
|
||||
persisted, either in the WAL (in the form of received precommit messages), or in the latest
|
||||
extended commit, the only way that vote extensions needed to start the next height could be lost
|
||||
forever would be if all validators crashed and never recovered (e.g. disk corruption).
|
||||
Since a *correct* node MUST eventually recover, this violates Tendermint's assumption of more than
|
||||
*2n<sub>h</sub>/3* correct validators for every height *h*.
|
||||
|
||||
No problem to solve here.
|
||||
|
||||
- **(c)** *Lagging majority*.
|
||||
|
||||
Let us assume the validator set does not change between *h* and *h+1*.
|
||||
It is not possible by the nature of the Tendermint algorithm, which requires more
|
||||
than *2n<sub>h</sub>/3* precommit votes for some round of height *h* in order to make progress.
|
||||
So, only up to *n<sub>h</sub>/3* validators can lag behind.
|
||||
|
||||
On the other hand, for the case where there are changes to the validator set between *h* and
|
||||
*h+1* please see case (d) below, where the extreme case is discussed.
|
||||
|
||||
- **(d)** *Validator set changes completely between* h *and* h+1.
|
||||
|
||||
If sets *valset<sub>h</sub>* and *valset<sub>h+1</sub>* are disjoint,
|
||||
more than *2n<sub>h</sub>/3* of validators in height *h* should
|
||||
have actively participated in conensus in *h*. So, as of height *h*, only a minority of validators
|
||||
in *h* can be lagging behind, although they could all lag behind from *h+1* on, as they are no
|
||||
longer validators, only full nodes. This situation falls under the assumptions of case (h) below.
|
||||
|
||||
As for validators in *valset<sub>h+1</sub>*, as they were not validators as of height *h*, they
|
||||
could all be lagging behind by that time. However, by the time *h* finishes and *h+1* begins, the
|
||||
chain will halt until more than *2n<sub>h+1</sub>/3* of them have caught up and started consensus
|
||||
at height *h+1*. If set *valset<sub>h+1</sub>* does not change in *h+2* and subsequent
|
||||
heights, only up to *n<sub>h+1</sub>/3* validators will be able to lag behind. Thus, we have
|
||||
converted this case into case (h) below.
|
||||
|
||||
- **(e)** *Enough validators crash to block the rest*.
|
||||
|
||||
In this case, blockchain progress halts, i.e. surviving full nodes keep increasing rounds
|
||||
indefinitely, until some of the crashed validators are able to recover.
|
||||
Those validators that recover first will handshake with the Application and recover at the height
|
||||
they crashed, which is still the same the nodes that did not crash are stuck in, so they don't need
|
||||
to catch up.
|
||||
Further, they had persisted the extended commit for the previous height. Nothing to solve.
|
||||
|
||||
For those validators recovering later, we are in case (h) below.
|
||||
|
||||
- **(f)** *Some validators crash, but not enough to block progress*.
|
||||
|
||||
When the correct processes that crashed recover, they handshake with the Application and resume at
|
||||
the height they were at when they crashed. As the blockchain did not stop making progress, the
|
||||
recovered processes are likely to have fallen behind with respect to the progressing majority.
|
||||
|
||||
At this point, the recovered processes are in case (h) below.
|
||||
|
||||
- **(g)** *A new full node starts*.
|
||||
|
||||
The reasoning here also applies to the case when more than one full node are starting.
|
||||
When the full node starts from scratch, it has no state (its current height is 0). Ignoring
|
||||
statesync for the time being, the node just needs to catch up by applying past blocks one by one
|
||||
(after verifying them).
|
||||
|
||||
Thus, the node is in case (h) below.
|
||||
|
||||
- **(h)** *Advancing majority, lagging minority*
|
||||
|
||||
In this case, some nodes are late. More precisely, at the present time, a set of full nodes,
|
||||
denoted *L<sub>h<sub>p</sub></sub>*, are falling behind
|
||||
(e.g., temporary disconnection or network partition, memory thrashing, crashes, new nodes)
|
||||
an arbitrary
|
||||
number of heights:
|
||||
between *h<sub>s</sub>* and *h<sub>p</sub>*, where *h<sub>s</sub> < h<sub>p</sub>*, and
|
||||
*h<sub>p</sub>* is the highest height
|
||||
any correct full node has reached so far.
|
||||
|
||||
The correct full nodes that reached *h<sub>p</sub>* were able to decide for *h<sub>p</sub>-1*.
|
||||
Therefore, less than *n<sub>h<sub>p</sub>-1</sub>/3* validators of *h<sub>p</sub>-1* can be part
|
||||
of *L<sub>h<sub>p</sub></sub>*, since enough up-to-date validators needed to actively participate
|
||||
in consensus for *h<sub>p</sub>-1*.
|
||||
|
||||
Since, at the present time,
|
||||
no node in *L<sub>h<sub>p</sub></sub>* took part in any consensus between
|
||||
*h<sub>s</sub>* and *h<sub>p</sub>-1*,
|
||||
the reasoning above can be extended to validator set changes between *h<sub>s</sub>* and
|
||||
*h<sub>p</sub>-1*. This results in the following restriction on the full nodes that can be part of *L<sub>h<sub>p</sub></sub>*.
|
||||
|
||||
- ∀ *h*, where *h<sub>s</sub> ≤ h < h<sub>p</sub>*,
|
||||
| *valset<sub>h</sub>* ∩ *L<sub>h<sub>p</sub></sub>* | *< n<sub>h</sub>/3*
|
||||
|
||||
If this property does not hold for a particular height *h*, where
|
||||
*h<sub>s</sub> ≤ h < h<sub>p</sub>*, Tendermint could not have progressed beyond *h* and
|
||||
therefore no full node could have reached *h<sub>p</sub>* (a contradiction).
|
||||
|
||||
These lagging nodes in *L<sub>h<sub>p</sub></sub>* need to catch up. They have to obtain the
|
||||
information needed to make
|
||||
progress from other nodes. For each height *h* between *h<sub>s</sub>* and *h<sub>p</sub>-2*,
|
||||
this includes the decided block for *h*, and the
|
||||
precommit votes also for *deciding h* (which can be extracted from the block at height *h+1*).
|
||||
|
||||
At a given height *h<sub>c</sub>* (where possibly *h<sub>c</sub> << h<sub>p</sub>*),
|
||||
a full node in *L<sub>h<sub>p</sub></sub>* will consider itself *caught up*, based on the
|
||||
(maybe out of date) information it is getting from its peers. Then, the node needs to be ready to
|
||||
propose at height *h<sub>c</sub>+1*, which requires having received the vote extensions for
|
||||
*h<sub>c</sub>*.
|
||||
As the vote extensions are *not* stored in the blocks, and it is difficult to have strong
|
||||
guarantees on *when* a late node considers itself caught up, providing the late node with the right
|
||||
vote extensions for the right height poses a problem.
|
||||
|
||||
At this point, we have described and compared all cases raised in discussions leading up to this
|
||||
RFC. The list above aims at being exhaustive. The analysis of each case included above makes all of
|
||||
them converge into case (h).
|
||||
|
||||
### Current Catch-up Mechanisms
|
||||
|
||||
We now briefly describe the current catch-up mechanisms in the reactors concerned in Tendermint.
|
||||
|
||||
#### Statesync
|
||||
|
||||
Full nodes optionally run statesync just after starting, when they start from scratch.
|
||||
If statesync succeeds, an Application snapshot is installed, and Tendermint jumps from height 0 directly
|
||||
to the height the Application snapshop represents, without applying the block of any previous height.
|
||||
Some light blocks are received and stored in the block store for running light-client verification of
|
||||
all the skipped blocks. Light blocks are incomplete blocks, typically containing the header and the
|
||||
canonical commit but, e.g., no transactions. They are stored in the block store as "signed headers".
|
||||
|
||||
The statesync reactor is not really relevant for solving the problem discussed in this RFC. We will
|
||||
nevertheless mention it when needed; in particular, to understand some corner cases.
|
||||
|
||||
#### Blocksync
|
||||
|
||||
The blocksync reactor kicks in after start up or recovery (and, optionally, after statesync is done)
|
||||
and sends the following messages to its peers:
|
||||
|
||||
- `StatusRequest` to query the height its peers are currently at, and
|
||||
- `BlockRequest`, asking for blocks of heights the local node is missing.
|
||||
|
||||
Using `BlockResponse` messages received from peers, the blocksync reactor validates each received
|
||||
block using the block of the following height, saves the block in the block store, and sends the
|
||||
block to the Application for execution.
|
||||
|
||||
If blocksync has validated and applied the block for the height *previous* to the highest seen in
|
||||
a `StatusResponse` message, or if no progress has been made after a timeout, the node considers
|
||||
itself as caught up and switches to the consensus reactor.
|
||||
|
||||
#### Consensus Reactor
|
||||
|
||||
The consensus reactor runs the full Tendermint algorithm. For a validator this means it has to
|
||||
propose blocks, and send/receive prevote/precommit messages, as mandated by Tendermint, before it can
|
||||
decide and move on to the next height.
|
||||
|
||||
If a full node that is running the consensus reactor falls behind at height *h*, when a peer node
|
||||
realises this it will retrieve the canonical commit of *h+1* from the block store, and *convert*
|
||||
it into a set of precommit votes and will send those to the late node.
|
||||
|
||||
## Discussion
|
||||
|
||||
### Solutions Proposed
|
||||
|
||||
These are the solutions proposed in discussions leading up to this RFC.
|
||||
|
||||
- **Solution 0.** *Vote extensions are made **best effort** in the specification*.
|
||||
|
||||
This is the simplest solution, considered as a way to provide vote extensions in a simple enough
|
||||
way so that it can be part of v0.36.
|
||||
It consists in changing the specification so as to not *require* that precommit votes used upon
|
||||
`PrepareProposal` contain their corresponding vote extensions. In other words, we render vote
|
||||
extensions optional.
|
||||
There are strong implications stemming from such a relaxation of the original specification.
|
||||
|
||||
- As a vote extension is signed *separately* from the vote it is extending, an intermediate node
|
||||
can now remove (i.e., censor) vote extensions from precommit messages at will.
|
||||
- Further, there is no point anymore in the spec requiring the Application to accept a vote extension
|
||||
passed via `VerifyVoteExtension` to consider a precommit message valid in its entirety. Remember
|
||||
this behavior of `VerifyVoteExtension` is adding a constraint to Tendermint's conditions for
|
||||
liveness.
|
||||
In this situation, it is better and simpler to just drop the vote extension rejected by the
|
||||
Application via `VerifyVoteExtension`, but still consider the precommit vote itself valid as long
|
||||
as its signature verifies.
|
||||
|
||||
- **Solution 1.** *Include vote extensions in the blockchain*.
|
||||
|
||||
Another obvious solution, which has somehow been considered in the past, is to include the vote
|
||||
extensions and their signatures in the blockchain.
|
||||
The blockchain would thus include the extended commit, rather than a regular commit, as the structure
|
||||
to be canonicalized in the next block.
|
||||
With this solution, the current mechanisms implemented both in the blocksync and consensus reactors
|
||||
would still be correct, as all the information a node needs to catch up, and to start proposing when
|
||||
it considers itself as caught-up, can now be recovered from past blocks saved in the block store.
|
||||
|
||||
This solution has two main drawbacks.
|
||||
|
||||
- As the block format must change, upgrading a chain requires a hard fork. Furthermore,
|
||||
all existing light client implementations will stop working until they are upgraded to deal with
|
||||
the new format (e.g., how certain hashes calculated and/or how certain signatures are checked).
|
||||
For instance, let us consider IBC, which relies on light clients. An IBC connection between
|
||||
two chains will be broken if only one chain upgrades.
|
||||
- The extra information (i.e., the vote extensions) that is now kept in the blockchain is not really
|
||||
needed *at every height* for a late node to catch up.
|
||||
- This information is only needed to be able to *propose* at the height the validator considers
|
||||
itself as caught-up. If a validator is indeed late for height *h*, it is useless (although
|
||||
correct) for it to call `PrepareProposal`, or `ExtendVote`, since the block is already decided.
|
||||
- Moreover, some use cases require pretty sizeable vote extensions, which would result in an
|
||||
important waste of space in the blockchain.
|
||||
|
||||
- **Solution 2.** *Skip* propose *step in Tendermint algorithm*.
|
||||
|
||||
This solution consists in modifying the Tendermint algorithm to skip the *send proposal* step in
|
||||
heights where the node does not have the required vote extensions to populate the call to
|
||||
`PrepareProposal`. The main idea behind this is that it should only happen when the validator is late
|
||||
and, therefore, up-to-date validators have already proposed (and decided) for that height.
|
||||
A small variation of this solution is, rather than skipping the *send proposal* step, the validator
|
||||
sends a special *empty* or *bottom* (⊥) proposal to signal other nodes that it is not ready to propose
|
||||
at (any round of) the current height.
|
||||
|
||||
The appeal of this solution is its simplicity. A possible implementation does not need to extend
|
||||
the data structures, or change the current catch-up mechanisms implemented in the blocksync or
|
||||
in the consensus reactor. When we lack the needed information (vote extensions), we simply rely
|
||||
on another correct validator to propose a valid block in other rounds of the current height.
|
||||
|
||||
However, this solution can be attacked by a byzantine node in the network in the following way.
|
||||
Let us consider the following scenario:
|
||||
|
||||
- all validators in *valset<sub>h</sub>* send out precommit messages, with vote extensions,
|
||||
for height *h*, round 0, roughly at the same time,
|
||||
- all those precommit messages contain non-`nil` precommit votes, which vote for block *b*
|
||||
- all those precommit messages sent in height *h*, round 0, and all messages sent in
|
||||
height *h*, round *r > 0* get delayed indefinitely, so,
|
||||
- all validators in *valset<sub>h</sub>* keep waiting for enough precommit
|
||||
messages for height *h*, round 0, needed for deciding in height *h*
|
||||
- an intermediate (malicious) full node *m* manages to receive block *b*, and gather more than
|
||||
*2n<sub>h</sub>/3* precommit messages for height *h*, round 0,
|
||||
- one way or another, the solution should have either (a) a mechanism for a full node to *tell*
|
||||
another full node it is late, or (b) a mechanism for a full node to conclude it is late based
|
||||
on other full nodes' messages; any of these mechanisms should, at the very least,
|
||||
require the late node receiving the decided block and a commit (not necessarily an extended
|
||||
commit) for *h*,
|
||||
- node *m* uses the gathered precommit messages to build a commit for height *h*, round 0,
|
||||
- in order to convince full nodes that they are late, node *m* either (a) *tells* them they
|
||||
are late, or (b) shows them it (i.e. *m*) is ahead, by sending them block *b*, along with the
|
||||
commit for height *h*, round 0,
|
||||
- all full nodes conclude they are late from *m*'s behavior, and use block *b* and the commit for
|
||||
height *h*, round 0, to decide on height *h*, and proceed to height *h+1*.
|
||||
|
||||
At this point, *all* full nodes, including all validators in *valset<sub>h+1</sub>*, have advanced
|
||||
to height *h+1* believing they are late, and so, expecting the *hypothetical* leading majority of
|
||||
validators in *valset<sub>h+1</sub>* to propose for *h+1*. As a result, the blockhain
|
||||
grinds to a halt.
|
||||
A (rather complex) ad-hoc mechanism would need to be carried out by node operators to roll
|
||||
back all validators to the precommit step of height *h*, round *r*, so that they can regenerate
|
||||
vote extensions (remember vote extensions are non-deterministic) and continue execution.
|
||||
|
||||
- **Solution 3.** *Require extended commits to be available at switching time*.
|
||||
|
||||
This one is more involved than all previous solutions, and builds on an idea present in Solution 2:
|
||||
vote extensions are actually not needed for Tendermint to make progress as long as the
|
||||
validator is *certain* it is late.
|
||||
|
||||
We define two modes. The first is denoted *catch-up mode*, and Tendermint only calls
|
||||
`FinalizeBlock` for each height when in this mode. The second is denoted *consensus mode*, in
|
||||
which the validator considers itself up to date and fully participates in consensus and calls
|
||||
`PrepareProposal`/`ProcessProposal`, `ExtendVote`, and `VerifyVoteExtension`, before calling
|
||||
`FinalizeBlock`.
|
||||
|
||||
The catch-up mode does not need vote extension information to make progress, as all it needs is the
|
||||
decided block at each height to call `FinalizeBlock` and keep the state-machine replication making
|
||||
progress. The consensus mode, on the other hand, does need vote extension information when
|
||||
starting every height.
|
||||
|
||||
Validators are in consensus mode by default. When a validator in consensus mode falls behind
|
||||
for whatever reason, e.g. cases (b), (d), (e), (f), (g), or (h) above, we introduce the following
|
||||
key safety property:
|
||||
|
||||
- for every height *h<sub>p</sub>*, a full node *f* in *h<sub>p</sub>* refuses to switch to catch-up
|
||||
mode **until** there exists a height *h'* such that:
|
||||
- *p* has received and (light-client) verified the blocks of
|
||||
all heights *h*, where *h<sub>p</sub> ≤ h ≤ h'*
|
||||
- it has received an extended commit for *h'* and has verified:
|
||||
- the precommit vote signatures in the extended commit
|
||||
- the vote extension signatures in the extended commit: each is signed with the same
|
||||
key as the precommit vote it extends
|
||||
|
||||
If the condition above holds for *h<sub>p</sub>*, namely receiving a valid sequence of blocks in
|
||||
the *f*'s future, and an extended commit corresponding to the last block in the sequence, then
|
||||
node *f*:
|
||||
|
||||
- switches to catch-up mode,
|
||||
- applies all blocks between *h<sub>p</sub>* and *h'* (calling `FinalizeBlock` only), and
|
||||
- switches back to consensus mode using the extended commit for *h'* to propose in the rounds of
|
||||
*h' + 1* where it is the proposer.
|
||||
|
||||
This mechanism, together with the invariant it uses, ensures that the node cannot be attacked by
|
||||
being fed a block without extensions to make it believe it is late, in a similar way as explained
|
||||
for Solution 2.
|
||||
|
||||
### Feasibility of the Proposed Solutions
|
||||
|
||||
Solution 0, besides the drawbacks described in the previous section, provides guarantees that are
|
||||
weaker than the rest. The Application does not have the assurance that more than *2n<sub>h</sub>/3* vote
|
||||
extensions will *always* be available when calling `PrepareProposal` at height *h+1*.
|
||||
This level of guarantees is probably not strong enough for vote extensions to be useful for some
|
||||
important use cases that motivated them in the first place, e.g., encrypted mempool transactions.
|
||||
|
||||
Solution 1, while being simple in that the changes needed in the current Tendermint codebase would
|
||||
be rather small, is changing the block format, and would therefore require all blockchains using
|
||||
Tendermint v0.35 or earlier to hard-fork when upgrading to v0.36.
|
||||
|
||||
Since Solution 2 can be attacked, one might prefer Solution 3, even if it is more involved
|
||||
to implement. Further, we must elaborate on how we can turn Solution 3, described in abstract
|
||||
terms in the previous section, into a concrete implementation compatible with the current
|
||||
Tendermint codebase.
|
||||
|
||||
### Current Limitations and Possible Implementations
|
||||
|
||||
The main limitations affecting the current version of Tendermint are the following.
|
||||
|
||||
- The current version of the blocksync reactor does not use the full
|
||||
[light client verification](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/light-client/README.md)
|
||||
algorithm to validate blocks coming from other peers.
|
||||
- The code being structured into the blocksync and consensus reactors, only switching from the
|
||||
blocksync reactor to the consensus reactor is supported; switching in the opposite direction is
|
||||
not supported. Alternatively, the consensus reactor could have a mechanism allowing a late node
|
||||
to catch up by skipping calls to `PrepareProposal`/`ProcessProposal`, and
|
||||
`ExtendVote`/`VerifyVoteExtension` and only calling `FinalizeBlock` for each height.
|
||||
Such a mechanism does not exist at the time of writing this RFC.
|
||||
|
||||
The blocksync reactor featuring light client verification is being actively worked on (tentatively
|
||||
for v0.37). So it is best if this RFC does not try to delve into that problem, but just makes sure
|
||||
its outcomes are compatible with that effort.
|
||||
|
||||
In subsection [Cases to Address](#cases-to-address), we concluded that we can focus on
|
||||
solving case (h) in theoretical terms.
|
||||
However, as the current Tendermint version does not yet support switching back to blocksync once a
|
||||
node has switched to consensus, we need to split case (h) into two cases. When a full node needs to
|
||||
catch up...
|
||||
|
||||
- **(h.1)** ... it has not switched yet from the blocksync reactor to the consensus reactor, or
|
||||
|
||||
- **(h.2)** ... it has already switched to the consensus reactor.
|
||||
|
||||
This is important in order to discuss the different possible implementations.
|
||||
|
||||
#### Base Implementation: Persist and Propagate Extended Commit History
|
||||
|
||||
In order to circumvent the fact that we cannot switch from the consensus reactor back to blocksync,
|
||||
rather than just keeping the few most recent extended commits, nodes will need to keep
|
||||
and gossip a backlog of extended commits so that the consensus reactor can still propose and decide
|
||||
in out-of-date heights (even if those proposals will be useless).
|
||||
|
||||
The base implementation - for which an experimental patch exists - consists in the conservative
|
||||
approach of persisting in the block store *all* extended commits for which we have also stored
|
||||
the full block. Currently, when statesync is run at startup, it saves light blocks.
|
||||
This base implementation does not seek
|
||||
to receive or persist extended commits for those light blocks as they would not be of any use.
|
||||
|
||||
Then, we modify the blocksync reactor so that peers *always* send requested full blocks together
|
||||
with the corresponding extended commit in the `BlockResponse` messages. This guarantees that the
|
||||
block store being reconstructed by blocksync has the same information as that of peers that are
|
||||
up to date (at least starting from the latest snapshot applied by statesync before starting blocksync).
|
||||
Thus, blocksync has all the data it requires to switch to the consensus reactor, as long as one of
|
||||
the following exit conditions are met:
|
||||
|
||||
- The node is still at height 0 (where no commit or extended commit is needed)
|
||||
- The node has processed at least 1 block in blocksync
|
||||
|
||||
The second condition is needed in case the node has installed an Application snapshot during statesync.
|
||||
If that is the case, at the time blocksync starts, the block store only has the data statesync has saved:
|
||||
light blocks, and no extended commits.
|
||||
Hence we need to blocksync at least one block from another node, which will be sent with its corresponding extended commit, before we can switch to consensus.
|
||||
|
||||
As a side note, a chain might be started at a height *h<sub>i</sub> > 0*, all other heights
|
||||
*h < h<sub>i</sub>* being non-existent. In this case, the chain is still considered to be at height 0 before
|
||||
block *h<sub>i</sub>* is applied, so the first condition above allows the node to switch to consensus even
|
||||
if blocksync has not processed any block (which is always the case if all nodes are starting from scratch).
|
||||
|
||||
When a validator falls behind while having already switched to the consensus reactor, a peer node can
|
||||
simply retrieve the extended commit for the required height from the block store and reconstruct a set of
|
||||
precommit votes together with their extensions and send them in the form of precommit messages to the
|
||||
validator falling behind, regardless of whether the peer node holds the extended commit because it
|
||||
actually participated in that consensus and thus received the precommit messages, or it received the extended commit via a `BlockResponse` message while running blocksync.
|
||||
|
||||
This solution requires a few changes to the consensus reactor:
|
||||
|
||||
- upon saving the block for a given height in the block store at decision time, save the
|
||||
corresponding extended commit as well
|
||||
- in the catch-up mechanism, when a node realizes that another peer is more than 2 heights
|
||||
behind, it uses the extended commit (rather than the canoncial commit as done previously) to
|
||||
reconstruct the precommit votes with their corresponding extensions
|
||||
|
||||
The changes to the blocksync reactor are more substantial:
|
||||
|
||||
- the `BlockResponse` message is extended to include the extended commit of the same height as
|
||||
the block included in the response (just as they are stored in the block store)
|
||||
- structure `bpRequester` is likewise extended to hold the received extended commits coming in
|
||||
`BlockResponse` messages
|
||||
- method `PeekTwoBlocks` is modified to also return the extended commit corresponding to the first block
|
||||
- when successfully verifying a received block, the reactor saves its corresponding extended commit in
|
||||
the block store
|
||||
|
||||
The two main drawbacks of this base implementation are:
|
||||
|
||||
- the increased size taken by the block store, in particular with big extensions
|
||||
- the increased bandwith taken by the new format of `BlockResponse`
|
||||
|
||||
#### Possible Optimization: Pruning the Extended Commit History
|
||||
|
||||
If we cannot switch from the consensus reactor back to the blocksync reactor we cannot prune the extended commit backlog in the block store without sacrificing the implementation's correctness. The asynchronous
|
||||
nature of our distributed system model allows a process to fall behing an arbitrary number of
|
||||
heights, and thus all extended commits need to be kept *just in case* a node that late had
|
||||
previously switched to the consensus reactor.
|
||||
|
||||
However, there is a possibility to optimize the base implementation. Every time we enter a new height,
|
||||
we could prune from the block store all extended commits that are more than *d* heights in the past.
|
||||
Then, we need to handle two new situations, roughly equivalent to cases (h.1) and (h.2) described above.
|
||||
|
||||
- (h.1) A node starts from scratch or recovers after a crash. In thisy case, we need to modify the
|
||||
blocksync reactor's base implementation.
|
||||
- when receiving a `BlockResponse` message, it MUST accept that the extended commit set to `nil`,
|
||||
- when sending a `BlockResponse` message, if the block store contains the extended commit for that
|
||||
height, it MUST set it in the message, otherwise it sets it to `nil`,
|
||||
- the exit conditions used for the base implementation are no longer valid; the only reliable exit
|
||||
condition now consists in making sure that the last block processed by blocksync was received with
|
||||
the corresponding commit, and not `nil`; this extended commit will allow the node to switch from
|
||||
the blocksync reactor to the consensus reactor and immediately act as a proposer if required.
|
||||
- (h.2) A node already running the consensus reactor falls behind beyond *d* heights. In principle,
|
||||
the node will be stuck forever as no other node can provide the vote extensions it needs to make
|
||||
progress (they all have pruned the corresponding extended commit).
|
||||
However we can manually have the node crash and recover as a workaround. This effectively converts
|
||||
this case into (h.1).
|
||||
|
||||
### Formalization Work
|
||||
|
||||
A formalization work to show or prove the correctness of the different use cases and solutions
|
||||
presented here (and any other that may be found) needs to be carried out.
|
||||
A question that needs a precise answer is how many extended commits (one?, two?) a node needs
|
||||
to keep in persistent memory when implementing Solution 3 described above without Tendermint's
|
||||
current limitations.
|
||||
Another important invariant we need to prove formally is that the set of vote extensions
|
||||
required to make progress will always be held somewhere in the network.
|
||||
|
||||
## References
|
||||
|
||||
- [ABCI++ specification](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/abci%2B%2B/README.md)
|
||||
- [ABCI as of v0.35](https://github.com/tendermint/spec/blob/4fb99af/spec/abci/README.md)
|
||||
- [Vote extensions issue](https://github.com/tendermint/tendermint/issues/8174)
|
||||
- [Light client verification](https://github.com/tendermint/tendermint/blob/4743a7ad0/spec/light-client/README.md)
|
||||
@@ -0,0 +1,555 @@
|
||||
# RFC 018: BLS Signature Aggregation Exploration
|
||||
|
||||
## Changelog
|
||||
|
||||
- 01-April-2022: Initial draft (@williambanfield).
|
||||
- 15-April-2022: Draft complete (@williambanfield).
|
||||
|
||||
## Abstract
|
||||
|
||||
## Background
|
||||
|
||||
### Glossary
|
||||
|
||||
The terms that are attached to these types of cryptographic signing systems
|
||||
become confusing quickly. Different sources appear to use slightly different
|
||||
meanings of each term and this can certainly add to the confusion. Below is
|
||||
a brief glossary that may be helpful in understanding the discussion that follows.
|
||||
|
||||
* **Short Signature**: A signature that does not vary in length with the
|
||||
number of signers.
|
||||
* **Multi-Signature**: A signature generated over a single message
|
||||
where, given the message and signature, a verifier is able to determine that
|
||||
all parties signed the message. May be short or may vary with the number of signers.
|
||||
* **Aggregated Signature**: A _short_ signature generated over messages with
|
||||
possibly different content where, given the messages and signature, a verifier
|
||||
should be able to determine that all the parties signed the designated messages.
|
||||
* **Threshold Signature**: A _short_ signature generated from multiple signers
|
||||
where, given a message and the signature, a verifier is able to determine that
|
||||
a large enough share of the parties signed the message. The identities of the
|
||||
parties that contributed to the signature are not revealed.
|
||||
* **BLS Signature**: An elliptic-curve pairing-based signature system that
|
||||
has some nice properties for short multi-signatures. May stand for
|
||||
*Boneh-Lynn-Schacham* or *Barreto-Lynn-Scott* depending on the context. A
|
||||
BLS signature is type of signature scheme that is distinct from other forms
|
||||
of elliptic-curve signatures such as ECDSA and EdDSA.
|
||||
* **Interactive**: Cryptographic scheme where parties need to perform one or
|
||||
more request-response cycles to produce the cryptographic material. For
|
||||
example, an interactive signature scheme may require the signer and the
|
||||
verifier to cooperate to create and/or verify the signature, rather than a
|
||||
signature being created ahead of time.
|
||||
* **Non-interactive**: Cryptographic scheme where parties do not need to
|
||||
perform any request-response cycles to produce the cryptographic material.
|
||||
|
||||
### Brief notes on pairing-based elliptic-curve cryptography
|
||||
|
||||
Pairing-based elliptic-curve cryptography is quite complex and relies on several
|
||||
types of high-level math. Cryptography, in general, relies on being able to find
|
||||
problems with an asymmetry between the difficulty of calculating the solution
|
||||
and verifying that a given solution is correct.
|
||||
|
||||
Pairing-based cryptography works by operating on mathematical functions that
|
||||
satisfy the property of **bilinear mapping**. This property is satisfied for
|
||||
functions `e` with values `P`, `Q`, `R` and `S` where `e(P, Q + R) = e(P, Q) * e(P, R)`
|
||||
and `e(P + S, Q) = e(P, Q) * e(S, Q)`. The most familiar example of this is
|
||||
exponentiation. Written in common notation, `g^P*(Q+R) = g^(P*Q) * g^(P*R)` for
|
||||
some value `g`.
|
||||
|
||||
Pairing-based elliptic-curve cryptography creates a bilinear mapping using
|
||||
elliptic curves over a finite field. With some original curve, you can define two groups,
|
||||
`G1` and `G2` which are points of the original curve _modulo_ different values.
|
||||
Finally, you define a third group `Gt`, where points from `G1` and `G2` satisfy
|
||||
the property of bilinearity with `Gt`. In this scheme, the function `e` takes
|
||||
as inputs points in `G1` and `G2` and outputs values in `Gt`. Succintly, given
|
||||
some point `P` in `G1` and some point `Q` in `G1`, `e(P, Q) = C` where `C` is in `Gt`.
|
||||
You can efficiently compute the mapping of points in `G1` and `G2` into `Gt`,
|
||||
but you cannot efficiently determine what points were summed and paired to
|
||||
produce the value in `Gt`.
|
||||
|
||||
Functions are then defined to map digital signatures, messages, and keys into
|
||||
and out of points of `G1` or `G2` and signature verification is the process
|
||||
of calculating if a set of values representing a message, public key, and digital
|
||||
signature produce the same value in `Gt` through `e`.
|
||||
|
||||
Signatures can be created as either points in `G1` with public keys being
|
||||
created as points in `G2` or vice versa. For the case of BLS12-381, the popular
|
||||
curve used, points in `G1` are represented with 48 bytes and points in `G2` are
|
||||
represented with 96 bytes. It is up to the implementer of the cryptosystem to
|
||||
decide which should be larger, the public keys or the signatures.
|
||||
|
||||
BLS signatures rely on pairing-based elliptic-curve cryptography to produce
|
||||
various types of signatures. For a more in-depth but still high level discussion
|
||||
pairing-based elliptic-curve cryptography, see Vitalik Buterin's post on
|
||||
[Exploring Elliptic Curve Pairings][vitalik-pairing-post]. For much more in
|
||||
depth discussion, see the specific paper on BLS12-381, [Short signatures from
|
||||
the Weil Pairing][bls-weil-pairing] and
|
||||
[Compact Multi-Signatures for Smaller Blockchains][multi-signatures-smaller-blockchains].
|
||||
|
||||
### Adoption
|
||||
|
||||
BLS signatures have already gained traction within several popular projects.
|
||||
|
||||
* Algorand is working on an implementation.
|
||||
* [Zcash][zcash-adoption] has adopted BLS12-381 into the protocol.
|
||||
* [Ethereum 2.0][eth-2-adoption] has adopted BLS12-381 into the protocol.
|
||||
* [Chia Network][chia-adoption] has adopted BLS for signing blocks.
|
||||
* [Ostracon][line-ostracon-pr], a fork of Tendermint has adopted BLS for signing blocks.
|
||||
|
||||
### What systems may be affected by adding aggregated signatures?
|
||||
|
||||
#### Gossip
|
||||
|
||||
Gossip could be updated to aggregate vote signatures during a consensus round.
|
||||
This appears to be of frankly little utility. Creating an aggregated signature
|
||||
incurs overhead, so frequently re-aggregating may incur a significant
|
||||
overhead. How costly this is is still subject to further investigation and
|
||||
performance testing.
|
||||
|
||||
Even if vote signatures were aggregated before gossip, each validator would still
|
||||
need to receive and verify vote extension data from each (individual) peer validator in
|
||||
order for consensus to proceed. That displaces any advantage gained by aggregating signatures across the vote message in the presence of vote extensions.
|
||||
|
||||
#### Block Creation
|
||||
|
||||
When creating a block, the proposer may create a small set of short
|
||||
multi-signatures and attach these to the block instead of including one
|
||||
signature per validator.
|
||||
|
||||
#### Block Verification
|
||||
|
||||
Currently, we verify each validator signature using the public key associated
|
||||
with that validator. With signature aggregation, verification of blocks would
|
||||
not verify many signatures individually, but would instead check the (single)
|
||||
multi-signature using the public keys stored by the validator. This would also
|
||||
require a mechanism for indicating which validators are included in the
|
||||
aggregated signature.
|
||||
|
||||
#### IBC Relaying
|
||||
|
||||
IBC would no longer need to transmit a large set of signatures when
|
||||
updating state. These state updates do not happen for every IBC packet, only
|
||||
when changing an IBC light client's view of the counterparty chain's state.
|
||||
General [IBC packets][ibc-packet] only contain enough information to correctly
|
||||
route the data to the counterparty chain.
|
||||
|
||||
IBC does persist commit signatures to the chain in these `MsgUpdateClient`
|
||||
message when updating state. This message would no longer need the full set
|
||||
of unique signatures and would instead only need one signature for all of the
|
||||
data in the header.
|
||||
|
||||
Adding BLS signatures would create a new signature type that must be
|
||||
understood by the IBC module and by the relayers. For some operations, such
|
||||
as state updates, the set of data written into the chain and received by the
|
||||
IBC module could be slightly smaller.
|
||||
|
||||
## Discussion
|
||||
|
||||
### What are the proposed benefits to aggregated signatures?
|
||||
|
||||
#### Reduce Block Size
|
||||
|
||||
At the moment, a commit contains a 64-byte (512-bit) signature for each validator
|
||||
that voted for the block. For the Cosmos Hub, which has 175 validators in the
|
||||
active set, this amounts to about 11 KiB per block. That gives an upper bound of
|
||||
around 113 GiB over the lifetime of the chain's 10.12M blocks. (Note, the Hub has
|
||||
increased the number of validators in the active set over time so the total
|
||||
signature size over the history of the chain is likely somewhat less than that).
|
||||
|
||||
Signature aggregation would only produce two signatures for the entire block.
|
||||
One for the yeas and one for the nays. Each BLS aggregated signature is 48
|
||||
bytes, per the [IETF standard of BLS signatures][bls-ietf-ecdsa-compare].
|
||||
Over the lifetime of the same Cosmos Hub chain, that would amount to about 1
|
||||
GB, a savings of 112 GB. While that is a large factor of reduction it's worth
|
||||
bearing in mind that, at [GCP's cost][gcp-storage-pricing] of $.026 USD per GB,
|
||||
that is a total savings of around $2.50 per month.
|
||||
|
||||
#### Reduce Signature Creation and Verification Time
|
||||
|
||||
From the [IETF draft standard on BLS Signatures][bls-ietf], BLS signatures can be
|
||||
created in 370 microseconds and verified in 2700 microseconds. Our current
|
||||
[Ed25519 implementation][voi-ed25519-perf] was benchmarked locally to take
|
||||
13.9 microseconds to produce a signature and 2.03 milliseconds to batch verify
|
||||
128 signatures, which is slightly fewer than the 175 in the Hub. blst, a popular
|
||||
implementation of BLS signature aggregation was benchmarked to perform verification
|
||||
on 100 signatures in 1.5 milliseconds [when run locally][blst-verify-bench]
|
||||
on an 8 thread machine and pre-aggregated public keys. It is worth noting that
|
||||
the `ed25519` library verification time grew steadily with the number of signatures,
|
||||
whereas the bls library verification time remains constant. This is because the
|
||||
number of operations used to verify a signature does not grow at all with the
|
||||
number of signatures included in the aggregate signature (as long as the signers
|
||||
signed over the same message data as is the case in Tendermint).
|
||||
|
||||
It is worth noting that this would also represent a _degredation_ in signature
|
||||
verification time for chains with small validator sets. When batch verifying
|
||||
only 32 signatures, our ed25519 library takes .57 milliseconds, whereas BLS
|
||||
would still require the same 1.5 milliseconds.
|
||||
|
||||
For massive validator sets, blst dominates, taking the same 1.5 milliseconds to
|
||||
check an aggregated signature from 1024 validators versus our ed25519 library's
|
||||
13.066 milliseconds to batch verify a set of that size.
|
||||
|
||||
#### Reduce Light-Client Verification Time
|
||||
|
||||
The light client aims to be a faster and lighter-weight way to verify that a
|
||||
block was voted on by a Tendermint network. The light client fetches
|
||||
Tendermint block headers and commit signatures, performing public key
|
||||
verification to ensure that the associated validator set signed the block.
|
||||
Reducing the size of the commit signature would allow the light client to fetch
|
||||
block data more quickly.
|
||||
|
||||
Additionally, the faster signature verification times of BLS signatures mean
|
||||
that light client verification would proceed more quickly.
|
||||
|
||||
However, verification of an aggregated signature is all-or-nothing. The verifier
|
||||
cannot check that some singular signer had a signature included in the block.
|
||||
Instead, the verifier must use all public keys to check if some signature
|
||||
was included. This does mean that any light client implementation must always
|
||||
be able to fetch all public keys for any height instead of potentially being
|
||||
able to check if some singular validator's key signed the block.
|
||||
|
||||
#### Reduce Gossip Bandwidth
|
||||
|
||||
##### Vote Gossip
|
||||
|
||||
It is possible to aggregate subsets of signatures during voting, so that the
|
||||
network need not gossip all *n* validator signatures to all *n* validators.
|
||||
Theoretically, subsets of the signatures could be aggregated during consensus
|
||||
and vote messages could carry those aggregated signatures. Implementing this
|
||||
would certainly increase the complexity of the gossip layer but could possibly
|
||||
reduce the total number of signatures required to be verified by each validator.
|
||||
|
||||
##### Block Gossip
|
||||
|
||||
A reduction in the block size as a result of signature aggregation would
|
||||
naturally lead to a reduction in the bandwidth required to gossip a block.
|
||||
Each validator would only send and receive the smaller aggregated signatures
|
||||
instead of the full list of multi-signatures as we have them now.
|
||||
|
||||
### What are the drawbacks to aggregated signatures?
|
||||
|
||||
#### Heterogeneous key types cannot be aggregated
|
||||
|
||||
Aggregation requires a specific signature algorithm, and our legacy signing schemes
|
||||
cannot be aggregated. In practice, this means that aggregated signatures could
|
||||
be created for a subset of validators using BLS signatures, and validators
|
||||
with other key types (such as Ed25519) would still have to be be separately
|
||||
propagated in blocks and votes.
|
||||
|
||||
#### Many HSMs do not support aggregated signatures
|
||||
|
||||
**Hardware Signing Modules** (HSM) are a popular way to manage private keys.
|
||||
They provide additional security for key management and should be used when
|
||||
possible for storing highly sensitive private key material.
|
||||
|
||||
Below is a list of popular HSMs along with their support for BLS signatures.
|
||||
|
||||
* YubiKey
|
||||
* [No support][yubi-key-bls-support]
|
||||
* Amazon Cloud HSM
|
||||
* [No support][cloud-hsm-support]
|
||||
* Ledger
|
||||
* [Lists support for the BLS12-381 curve][ledger-bls-announce]
|
||||
|
||||
I cannot find support listed for Google Cloud, although perhaps it exists.
|
||||
|
||||
## Feasibility of implementation
|
||||
|
||||
This section outlines the various hurdles that would exist to implementing BLS
|
||||
signature aggregation into Tendermint. It aims to demonstrate that we _could_
|
||||
implement BLS signatures but that it would incur risk and require breaking changes for a
|
||||
reasonably unclear benefit.
|
||||
|
||||
### Can aggregated signatures be added as soft-upgrades?
|
||||
|
||||
In my estimation, yes. With the implementation of proposer-based timestamps,
|
||||
all validators now produce signatures on only one of two messages:
|
||||
|
||||
1. A [CanonicalVote][canonical-vote-proto] where the BlockID is the hash of the block or
|
||||
2. A `CanonicalVote` where the `BlockID` is nil.
|
||||
|
||||
The block structure can be updated to perform hashing and validation in a new
|
||||
way as a soft upgrade. This would look like adding a new section to the [Block.Commit][commit-proto] structure
|
||||
alongside the current `Commit.Signatures` field. This new field, tentatively named
|
||||
`AggregatedSignature` would contain the following structure:
|
||||
|
||||
```proto
|
||||
message AggregatedSignature {
|
||||
// yeas is a BitArray representing which validators in the active validator
|
||||
// set issued a 'yea' vote for the block.
|
||||
tendermint.libs.bits.BitArray yeas = 1;
|
||||
|
||||
// absent is a BitArray representing which validators in the active
|
||||
// validator set did not issue votes for the block.
|
||||
tendermint.libs.bits.BitArray absent = 2;
|
||||
|
||||
// yea_signature is an aggregated signature produced from all of the vote
|
||||
// signatures for the block.
|
||||
repeated bytes yea_signature = 3;
|
||||
|
||||
// nay_signature is an aggregated signature produced from all of the vote
|
||||
// signatures from votes for 'nil' for this block.
|
||||
// nay_signature should be made from all of the validators that were both not
|
||||
// in the 'yeas' BitArray and not in the 'absent' BitArray.
|
||||
repeated bytes nay_signature = 4;
|
||||
}
|
||||
```
|
||||
|
||||
Adding this new field as a soft upgrade would mean hashing this data structure
|
||||
into the blockID along with the old `Commit.Signatures` when both are present
|
||||
as well as ensuring that the voting power represented in the new
|
||||
`AggregatedSignature` and `Signatures` field was enough to commit the block
|
||||
during block validation. One can certainly imagine other possible schemes for
|
||||
implementing this but the above should serve as a simple enough proof of concept.
|
||||
|
||||
### Implementing vote-time and commit-time signature aggregation separately
|
||||
|
||||
Implementing aggregated BLS signatures as part of the block structure can easily be
|
||||
achieved without implementing any 'vote-time' signature aggregation.
|
||||
The block proposer would gather all of the votes, complete with signatures,
|
||||
as it does now, and produce a set of aggregate signatures from all of the
|
||||
individual vote signatures.
|
||||
|
||||
Implementing 'vote-time' signature aggregation cannot be achieved without
|
||||
also implementing commit-time signature aggregation. This is because such
|
||||
signatures cannot be dis-aggregated into their constituent pieces. Therefore,
|
||||
in order to implement 'vote-time' signature aggregation, we would need to
|
||||
either first implement 'commit-time' signature aggregation, or implement both
|
||||
'vote-time' signature aggregation while also updating the block creation and
|
||||
verification protocols to allow for aggregated signatures.
|
||||
|
||||
### Updating IBC clients
|
||||
|
||||
In order for IBC clients to function, they must be able to perform light-client
|
||||
verification of blocks on counterparty chains. Because BLS signatures are not
|
||||
currently part of light-clients, chains that transmit messages over IBC
|
||||
cannot update to using BLS signatures without their counterparties first
|
||||
being upgraded to parse and verify BLS. If chains upgrade without their
|
||||
counterparties first updating, they will lose the ability to interoperate with
|
||||
non-updated chains.
|
||||
|
||||
### New attack surfaces
|
||||
|
||||
BLS signatures and signature aggregation comes with a new set of attack surfaces.
|
||||
Additionally, it's not clear that all possible major attacks are currently known
|
||||
on the BLS aggregation schemes since new ones have been discovered since the ietf
|
||||
draft standard was written. The known attacks are manageable and are listed below.
|
||||
Our implementation would need to prevent against these but this does not appear
|
||||
to present a significant hurdle to implementation.
|
||||
|
||||
#### Rogue key attack prevention
|
||||
|
||||
Generating an aggregated signature requires guarding against what is called
|
||||
a [rogue key attack][bls-ietf-terms]. A rogue key attack is one in which a
|
||||
malicious actor can craft an _aggregate_ key that can produce signatures that
|
||||
appear to include a signature from a private key that the malicious actor
|
||||
does not actually know. In Tendermint terms, this would look like a Validator
|
||||
producing a vote signed by both itself and some other validator where the other
|
||||
validator did not actually produce the vote itself.
|
||||
|
||||
The main mechanisms for preventing this require that each entity prove that it
|
||||
can can sign data with just their private key. The options involve either
|
||||
ensuring that each entity sign a _different_ message when producing every
|
||||
signature _or_ producing a [proof of possession][bls-ietf-pop] (PoP) when announcing
|
||||
their key to the network.
|
||||
|
||||
A PoP is a message that demonstrates ownership of a private
|
||||
key. A simple scheme for PoP is one where the entity announcing
|
||||
its new public key to the network includes a digital signature over the bytes
|
||||
of the public key generated using the associated private key. Everyone receiving
|
||||
the public key and associated proof-of-possession can easily verify the
|
||||
signature and be sure the entity owns the private key.
|
||||
|
||||
This PoP scheme suits the Tendermint use case quite well since
|
||||
validator keys change infrequently so the associated PoPs would not be onerous
|
||||
to produce, verify, and store. Using this scheme allows signature verification
|
||||
to proceed more quickly, since all signatures are over identical data and
|
||||
can therefore be checked using an aggregated public key instead of one at a
|
||||
time, public key by public key.
|
||||
|
||||
#### Summing Zero Attacks
|
||||
|
||||
[Summing zero attacks][summing-zero-paper] are attacks that rely on using the '0' point of an
|
||||
elliptic curve. For BLS signatures, if the point 0 is chosen as the private
|
||||
key, then the 0 point will also always be the public key and all signatures
|
||||
produced by the key will also be the 0 point. This is easy enough to
|
||||
detect when verifying each signature individually.
|
||||
|
||||
However, because BLS signature aggregation creates an aggregated signature and
|
||||
an aggregated public key, a set of colluding signers can create a pair or set
|
||||
of signatures that are non-zero but which aggregate ("sum") to 0. The signatures that sum zero along with the
|
||||
summed public key of the colluding signers will verify any message. This would
|
||||
allow the colluding signers to sign any block or message with the same signature.
|
||||
This would be reasonably easy to detect and create evidence for because, in
|
||||
all other cases, the same signature should not verify more than message. It's
|
||||
not exactly clear how such an attack would advantage the colluding validators
|
||||
because the normal mechanisms of evidence gathering would still detect the
|
||||
double signing, regardless of the signatures on both blocks being identical.
|
||||
|
||||
### Backwards Compatibility
|
||||
|
||||
Backwards compatibility is an important consideration for signature verification.
|
||||
Specifically, it is important to consider whether chains using current versions
|
||||
of IBC would be able to interact with chains adopting BLS.
|
||||
|
||||
Because the `Block` shared by IBC and Tendermint is produced and parsed using
|
||||
protobuf, new structures can be added to the Block without breaking the
|
||||
ability of legacy users to parse the new structure. Breaking changes between
|
||||
current users of IBC and new Tendermint blocks only occur if data that is
|
||||
relied upon by the current users is no longer included in the current fields.
|
||||
|
||||
For the case of BLS aggregated signatures, a new `AggregatedSignature` field
|
||||
can therefore be added to the `Commit` field without breaking current users.
|
||||
Current users will be broken when counterparty chains upgrade to the new version
|
||||
and _begin using_ BLS signatures. Once counterparty chains begin using BLS
|
||||
signatures, the BlockID hashes will include hashes of the `AggregatedSignature`
|
||||
data structure that the legacy users will not be able to compute. Additionally,
|
||||
the legacy software will not be able to parse and verify the signatures to
|
||||
ensure that a supermajority of validators from the counterparty chain signed
|
||||
the block.
|
||||
|
||||
### Library Support
|
||||
|
||||
Libraries for BLS signature creation are limited in number, although active
|
||||
development appears to be ongoing. Cryptographic algorithms are difficult to
|
||||
implement correctly and correctness issues are extremely serious and dangerous.
|
||||
No further exploration of BLS should be undertaken without strong assurance of
|
||||
a well-tested library with continuing support for creating and verifying BLS
|
||||
signatures.
|
||||
|
||||
At the moment, there is one candidate, `blst`, that appears to be the most
|
||||
mature and well vetted. While this library is undergoing continuing auditing
|
||||
and is supported by funds from the Ethereum foundation, adopting a new cryptographic
|
||||
library presents some serious risks. Namely, if the support for the library were
|
||||
to be discontinued, Tendermint may become saddled with the requirement of supporting
|
||||
a very complex piece of software or force a massive ecosystem-wide migration away
|
||||
from BLS signatures.
|
||||
|
||||
This is one of the more serious reasons to avoid adopting BLS signatures at this
|
||||
time. There is no gold standard library. Some projects look promising, but no
|
||||
project has been formally verified with a long term promise of being supported
|
||||
well into the future.
|
||||
|
||||
#### Go Standard Library
|
||||
|
||||
The Go Standard library has no implementation of BLS signatures.
|
||||
|
||||
#### BLST
|
||||
|
||||
[blst][blst], or 'blast' is an implementation of BLS signatures written in C
|
||||
that provides bindings into Go as part of the repository. This library is
|
||||
actively undergoing formal verification by Galois and previously received an
|
||||
initial audit by NCC group, a firm I'd never heard of.
|
||||
|
||||
`blst` is [targeted for use in prysm][prysm-blst], the golang implementation of Ethereum 2.0.
|
||||
|
||||
#### Gnark-Crypto
|
||||
|
||||
[Gnark-Crypto][gnark] is a Go-native implementation of elliptic-curve pairing-based
|
||||
cryptography. It is not audited and is documented as 'as-is', although
|
||||
development appears to be active so formal verification may be forthcoming.
|
||||
|
||||
#### CIRCL
|
||||
|
||||
[CIRCL][circl] is a go-native implementation of several cryptographic primitives,
|
||||
bls12-381 among them. The library is written and maintained by Cloudflare and
|
||||
appears to receive frequent contributions. However, it lists itself as experimental
|
||||
and urges users to take caution before using it in production.
|
||||
|
||||
### Added complexity to light client verification
|
||||
|
||||
Implementing BLS signature aggregation in Tendermint would pose issues for the
|
||||
light client. The light client currently validates a subset of the signatures
|
||||
on a block when performing the verification algorithm. This is no longer possible
|
||||
with an aggregated signature. Aggregated signature verification is all-or-nothing.
|
||||
The light client could no longer check that a subset of validators from some
|
||||
set of validators is represented in the signature. Instead, it would need to create
|
||||
a new aggregated key with all the stated signers for each height it verified where
|
||||
the validator set changed.
|
||||
|
||||
This means that the speed advantages gained by using BLS cannot be fully realized
|
||||
by the light client since the client needs to perform the expensive operation
|
||||
of re-aggregating the public key. Aggregation is _not_ constant time in the
|
||||
number of keys and instead grows linearly. When [benchmarked locally][blst-verify-bench-agg],
|
||||
blst public key aggregation of 128 keys took 2.43 milliseconds. This, along with
|
||||
the 1.5 milliseconds to verify a signature would raise light client signature
|
||||
verification time to 3.9 milliseconds, a time above the previously mentioned
|
||||
batch verification time using our ed25519 library of 2.0 milliseconds.
|
||||
|
||||
Schemes to cache aggregated subsets of keys could certainly cut this time down at the
|
||||
cost of adding complexity to the light client.
|
||||
|
||||
### Added complexity to evidence handling
|
||||
|
||||
Implementing BLS signature aggregation in Tendermint would add complexity to
|
||||
the evidence handling within Tendermint. Currently, the light client can submit
|
||||
evidence of a fork attempt to the chain. This evidence consists of the set of
|
||||
validators that double-signed, including their public keys, with the conflicting
|
||||
block.
|
||||
|
||||
We can quickly check that the listed validators double signed by verifying
|
||||
that each of their signatures are in the submitted conflicting block. A BLS
|
||||
signature scheme would change this by requiring the light client to submit
|
||||
the public keys of all of the validators that signed the conflicting block so
|
||||
that the aggregated signature may be checked against the full signature set.
|
||||
Again, aggregated signature verification is all-or-nothing, so without all of
|
||||
the public keys, we cannot verify the signature at all. These keys would be
|
||||
retrievable. Any party that wanted to create a fork would want to convince a
|
||||
network that its fork is legitimate, so it would need to gossip the public keys.
|
||||
This does not hamper the feasibility of implementing BLS signature aggregation
|
||||
into Tendermint, but does represent yet another piece of added complexity to
|
||||
the associated protocols.
|
||||
|
||||
## Open Questions
|
||||
|
||||
* *Q*: Can you aggregate Ed25519 signatures in Tendermint?
|
||||
* There is a suggested scheme in github issue [7892][suggested-ed25519-agg],
|
||||
but additional rigor would be required to fully verify its correctness.
|
||||
|
||||
## Current Consideration
|
||||
|
||||
Adopting a signature aggregation scheme presents some serious risks and costs
|
||||
to the Tendermint project. It requires multiple backwards-incompatible changes
|
||||
to the code, namely a change in the structure of the block and a new backwards-incompatible
|
||||
signature and key type. It risks adding a new signature type for which new attack
|
||||
types are still being discovered _and_ for which no industry standard, battle-tested
|
||||
library yet exists.
|
||||
|
||||
The gains boasted by this new signing scheme are modest: Verification time is
|
||||
marginally faster and block sizes shrink by a few kilobytes. These are relatively
|
||||
minor gains in exchange for the complexity of the change and the listed risks of the technology.
|
||||
We should take a wait-and-see approach to BLS signature aggregation, monitoring
|
||||
the up-and-coming projects and consider implementing it as the libraries and
|
||||
standards develop.
|
||||
|
||||
### References
|
||||
|
||||
[line-ostracon-repo]: https://github.com/line/ostracon
|
||||
[line-ostracon-pr]: https://github.com/line/ostracon/pull/117
|
||||
[mit-BLS-lecture]: https://youtu.be/BFwc2XA8rSk?t=2521
|
||||
[gcp-storage-pricing]: https://cloud.google.com/storage/pricing#north-america_2
|
||||
[yubi-key-bls-support]: https://github.com/Yubico/yubihsm-shell/issues/66
|
||||
[cloud-hsm-support]: https://docs.aws.amazon.com/cloudhsm/latest/userguide/pkcs11-key-types.html
|
||||
[bls-ietf]: https://datatracker.ietf.org/doc/html/draft-irtf-cfrg-bls-signature-04
|
||||
[bls-ietf-terms]: https://datatracker.ietf.org/doc/html/draft-irtf-cfrg-bls-signature-04#section-1.3
|
||||
[bls-ietf-pop]: https://datatracker.ietf.org/doc/html/draft-irtf-cfrg-bls-signature-04#section-3.3
|
||||
[multi-signatures-smaller-blockchains]: https://eprint.iacr.org/2018/483.pdf
|
||||
[ibc-tendermint]: https://github.com/cosmos/ibc/tree/master/spec/client/ics-007-tendermint-client
|
||||
[zcash-adoption]: https://github.com/zcash/zcash/issues/2502
|
||||
[chia-adoption]: https://github.com/Chia-Network/chia-blockchain#chia-blockchain
|
||||
[bls-ietf-ecdsa-compare]: https://datatracker.ietf.org/doc/html/draft-irtf-cfrg-bls-signature-04#section-1.1
|
||||
[voi-ed25519-perf]: https://github.com/williambanfield/curve25519-voi/blob/benchmark/primitives/ed25519/PERFORMANCE.txt#L79
|
||||
[blst-verify-bench]: https://github.com/williambanfield/blst/blame/bench/bindings/go/PERFORMANCE.md#L9
|
||||
[blst-verify-bench-agg]: https://github.com/williambanfield/blst/blame/bench/bindings/go/PERFORMANCE.md#L23
|
||||
[vitalik-pairing-post]: https://medium.com/@VitalikButerin/exploring-elliptic-curve-pairings-c73c1864e627
|
||||
[ledger-bls-announce]: https://www.ledger.com/first-ever-firmware-update-coming-to-the-ledger-nano-x
|
||||
[commit-proto]: https://github.com/tendermint/tendermint/blob/be7cb50bb3432ee652f88a443e8ee7b8ef7122bc/proto/tendermint/types/types.proto#L121
|
||||
[canonical-vote-proto]: https://github.com/tendermint/tendermint/blob/be7cb50bb3432ee652f88a443e8ee7b8ef7122bc/spec/core/encoding.md#L283
|
||||
[blst]: https://github.com/supranational/blst
|
||||
[prysm-blst]: https://github.com/prysmaticlabs/prysm/blob/develop/go.mod#L75
|
||||
[gnark]: https://github.com/ConsenSys/gnark-crypto/
|
||||
[eth-2-adoption]: https://notes.ethereum.org/@GW1ZUbNKR5iRjjKYx6_dJQ/Skxf3tNcg_
|
||||
[bls-weil-pairing]: https://www.iacr.org/archive/asiacrypt2001/22480516.pdf
|
||||
[summing-zero-paper]: https://eprint.iacr.org/2021/323.pdf
|
||||
[circl]: https://github.com/cloudflare/circl
|
||||
[light-client-evidence]: https://github.com/tendermint/tendermint/blob/a6fd1fe20116d4b1f7e819cded81cece8e5c1ac7/types/evidence.go#L245
|
||||
[suggested-ed25519-agg]: https://github.com/tendermint/tendermint/issues/7892
|
||||
@@ -0,0 +1,400 @@
|
||||
# RFC 019: Configuration File Versioning
|
||||
|
||||
## Changelog
|
||||
|
||||
- 19-Apr-2022: Initial draft (@creachadair)
|
||||
- 20-Apr-2022: Updates from review feedback (@creachadair)
|
||||
|
||||
## Abstract
|
||||
|
||||
Updating configuration settings is an essential part of upgrading an existing
|
||||
node to a new version of the Tendermint software. Unfortunately, it is also
|
||||
currently a very manual process. This document discusses some of the history of
|
||||
changes to the config format, actions we've taken to improve the tooling for
|
||||
configuration upgrades, and additional steps we may want to consider.
|
||||
|
||||
## Background
|
||||
|
||||
A Tendermint node reads configuration settings at startup from a TOML formatted
|
||||
text file, typically named `config.toml`. The contents of this file are defined
|
||||
by the [`github.com/tendermint/tendermint/config`][config-pkg].
|
||||
|
||||
Although many settings in this file remain valid from one version of Tendermint
|
||||
to the next, new versions of Tendermint often add, update, and remove settings.
|
||||
These changes often require manual intervention by operators who are upgrading
|
||||
their nodes.
|
||||
|
||||
I propose we should provide better tools and documentation to help operators
|
||||
make configuration changes correctly during version upgrades. Ideally, as much
|
||||
as possible of any configuration file update should be automated, and where
|
||||
that is not possible or practical, we should provide clear, explicit directions
|
||||
for what steps need to be taken manually. Moreover, when the node discovers
|
||||
incorrect or invalid configuration, we should improve the diagnostics it emits
|
||||
so that the operator can quickly and easily find the relevant documentation,
|
||||
without having to grep through source code.
|
||||
|
||||
## Discussion
|
||||
|
||||
By convention, we are supposed to document required changes to the config file
|
||||
in the `UPGRADING.md` file for the release that introduces them. Although we
|
||||
have mostly done this, the level of detail in the upgrading instructions is
|
||||
often insufficient for an operator to correctly update their file.
|
||||
|
||||
The updates vary widely in complexity: Operators may need to add new required
|
||||
settings, update obsolete values for existing settings, move or rename existing
|
||||
settings within the file, or remove obsolete settings (which are thus invalid).
|
||||
Here are a few examples of each of these cases:
|
||||
|
||||
- **New required settings:** Tendermint v0.35 added a new top-level `mode`
|
||||
setting that determines whether a node runs as a validator, a full node, or a
|
||||
seed node. The default value is `"full"`, which means the operator of a
|
||||
validator must manually add `mode = "validator"` (or set the `--mode` flag on
|
||||
the command line) for their node to come up in the correct mode.
|
||||
|
||||
- **Updated obsolete values:** Tendermint v0.35 removed support for versions
|
||||
`"v1"` and `"v2"` of the blocksync (formerly "fastsync") protocol, requiring
|
||||
any node using either of those values to update to `"v0"`.
|
||||
|
||||
- **Moved/renamed settings:** Version v0.34 moved the top-level `pprof_laddr`
|
||||
setting under the `[rpc]` section.
|
||||
|
||||
Version v0.35 renamed every setting in the file from `snake_case` to
|
||||
`kebab-case`, moved the top-level `fast_sync` setting into the `[blocksync]`
|
||||
section as (itself renamed from `[fastsync]`), and moved all the top-level
|
||||
`priv-validator-*` settings under a new `[priv-validator]` section with their
|
||||
prefix trimmed off.
|
||||
|
||||
- **Removed obsolete settings:** Version v0.34 removed the `index_all_keys` and
|
||||
`index_keys` settings from the `[tx_index]` section; version v0.35 removed
|
||||
the `wal-dir` setting from the `[mempool]` section, and version v0.36 removed
|
||||
the `[blocksync]` section entirely.
|
||||
|
||||
While many of these changes are mentioned in the config section of the upgrade
|
||||
instructions, some are not mentioned at all, or are hidden in other parts of
|
||||
the doc. For instance, the v0.34 `pprof_laddr` change was documented only as an
|
||||
RPC flag change. (A savvy reader might realize that the flag `--rpc.pprof_laddr`
|
||||
implies a corresponding config section, but it omits the related detail that
|
||||
there was a top-level setting that's been renamed). The lesson here is not
|
||||
that the docs are bad, but to point out that prose is not the most efficient
|
||||
format to convey detailed changes like this. The upgrading instructions are
|
||||
still valuable for the human reader to understand what to expect.
|
||||
|
||||
### Concrete Steps
|
||||
|
||||
As part of the v0.36 development cycle, we spent some time reverse-engineering
|
||||
the configuration changes since the v0.34 release and built an experimental
|
||||
command-line tool called [`confix`][confix], whose job it is to automatically
|
||||
update the settings in a `config.toml` file to the latest version. We also
|
||||
backported a version of this tool into the v0.35.x branch at release v0.35.4.
|
||||
|
||||
This tool should work fine for configuration files created by Tendermint v0.34
|
||||
and later, but does not (yet) know how to handle changes from prior versions of
|
||||
Tendermint. Part of the difficulty for older versions is simply logistical: To
|
||||
figure out which changes to apply, we need to understand something about the
|
||||
version that made the file, as well as the version we're converting it to.
|
||||
|
||||
> **Discussion point:** In the future we might want to consider incorporating
|
||||
> this into the node CLI directly, but we're keeping it separate for now until
|
||||
> we can get some feedback from operators.
|
||||
|
||||
For the experiment, we handled this by carefully searching the history of
|
||||
config format changes for shibboleths to bound the version: For example, the
|
||||
`[fastsync]` section was added in Tendermint v0.32 and renamed `[blocksync]` in
|
||||
Tendermint v0.35. So if we see a `[fastsync]` section, we have some confidence
|
||||
that the file was created by v0.32, v0.33, or v0.34.
|
||||
|
||||
But such signals are delicate: The `[blocksync]` section was removed in v0.36,
|
||||
so if we do not find `[fastsync]`, we cannot conclude from that alone that the
|
||||
file is from v0.31 or earlier -- we have to look for corroborating details.
|
||||
While such "sniffing" tactics are fine for an experiment, they aren't as robust
|
||||
as we might like.
|
||||
|
||||
This is especially relevant for configuration files that may have already been
|
||||
manually upgraded across several versions by the time we are asked to update
|
||||
them again. Another related concern is that we'd like to make sure conversion
|
||||
is idempotent, so that it would be safe to rerun the tool over an
|
||||
already-converted file without breaking anything.
|
||||
|
||||
### Config Versioning
|
||||
|
||||
One obvious tactic we could use for future releases is add a version marker to
|
||||
the config file. This would give tools like `confix` (and the node itself) a
|
||||
way to calibrate their expectations. Rather than being a version for the file
|
||||
itself, however, this version marker would indicate which version of Tendermint
|
||||
is needed to read the file.
|
||||
|
||||
Provisionally, this might look something like:
|
||||
|
||||
```toml
|
||||
# THe minimum version of Tendermint compatible with the contents of
|
||||
# this configuration file.
|
||||
config-version = 'v0.35'
|
||||
```
|
||||
|
||||
When initializing a new node, Tendermint would populate this field with its own
|
||||
version (e.g., `v0.36`). When conducting an upgrade, tools like `confix` can
|
||||
then use this to decide which conversions are valid, and then update the value
|
||||
accordingly. After converting a file marked `'v0.35'` to`'v0.37'`, the
|
||||
conversion tool sets the file's `config-version` to reflect its compatibility.
|
||||
|
||||
> **Discussion point:** This example presumes we would keep config files
|
||||
> compatible within a given release cycle, e.g., all of v0.36.x. We could also
|
||||
> use patch numbers here, if we think there's some reason to permit changes
|
||||
> that would require config file edits at that granularity. I don't think we
|
||||
> should, but that's a design question to consider.
|
||||
|
||||
Upon seeing an up-to-date version marker, the conversion tool can simply exit
|
||||
with a diagnostic like "this file is already up-to-date", rather than sniffing
|
||||
the keyspace and potentially introducing errors. In addition, this would let a
|
||||
tool detect config files that are _newer_ than the one it understands, and
|
||||
issue a safe diagnostic rather than doing something wrong. Plus, besides
|
||||
avoiding potentially unsafe conversions, this would also serve as
|
||||
human-readable documentation that the file is up-to-date for a given version.
|
||||
|
||||
Adding a config version would not address the problem of how to convert files
|
||||
created by older versions of Tendermint, but it would at least help us build
|
||||
more robust config tooling going forward.
|
||||
|
||||
### Stability and Change
|
||||
|
||||
In light of the discussion so far, it is natural to examine why we make so many
|
||||
changes to the configuration file from one version to the next, and whether we
|
||||
could reduce friction by being more conservative about what we make
|
||||
configurable, what config changes we make over time, and how we roll them out.
|
||||
|
||||
Some changes, like renaming everything from snake case to kebab case, are
|
||||
entirely gratuitous. We could safely agree not to make those kinds of changes.
|
||||
Apart from that obvious case, however, many other configuration settings
|
||||
provide value to node operators in cases where there is no simple, universal
|
||||
setting that matches every application.
|
||||
|
||||
Taking a high-level view, there are several broad reasons why we might want to
|
||||
make changes to configuration settings:
|
||||
|
||||
- **Lessons learned:** Configuration settings are a good way to try things out
|
||||
in production, before making more invasive changes to the consensus protocol.
|
||||
|
||||
For example, up until Tendermint v0.35, consensus timeouts were specified as
|
||||
per-node configuration settings (e.g., `timeout-precommit` et al.). This
|
||||
allowed operators to tune these values for the needs of their network, but
|
||||
had the downside that individually-misconfigured nodes could stall consensus.
|
||||
|
||||
Based on that experience, these timeouts have been deprecated in Tendermint
|
||||
v0.36 and converted to consensus parameters, to be consistent across all
|
||||
nodes in the network.
|
||||
|
||||
- **Migration & experimentation:** Introducing new features and updating old
|
||||
features can complicate migration for existing users of the software.
|
||||
Temporary or "experimental" configuration settings can be a valuable way to
|
||||
mitigate that friction.
|
||||
|
||||
For example, Tendermint v0.36 introduces a new RPC event subscription
|
||||
endpoint (see [ADR 075][adr075]) that will eventually replace the existing
|
||||
webwocket-based interface. To give users time to migrate, v0.36 adds an
|
||||
`experimental-disable-websocket` setting, defaulted to `false`, that allows
|
||||
operators to selectively disable the websocket API for testing purposes
|
||||
during the conversion. This setting is designed to be removed in v0.37, when
|
||||
the old interface is no longer supported.
|
||||
|
||||
- **Ongoing maintenance:** Sometimes configuration settings become obsolete,
|
||||
and the cost of removing them trades off against the potential risks of
|
||||
leaving a non-functional or deprecated knob hooked up indefinitely.
|
||||
|
||||
For example, Tendermint v0.35 deprecated two alternate implementations of the
|
||||
blocksync protocol, one of which was deleted entirely (`v1`) and one of which
|
||||
was scheduled for removal (`v2`). The `blocksync.version` setting, which had
|
||||
been added as a migration aid, became obsolete and needed to be updated.
|
||||
|
||||
Despite our best intentions, sometimes engineering designs do not work out.
|
||||
It's just as important to leave room to back out of changes we have since
|
||||
reconsidered, as it is to support migrations forward onto new and improved
|
||||
code.
|
||||
|
||||
- **Clarity and legibility:** Besides configuring the software, another
|
||||
important purpose of a config file is to document intent for the humans who
|
||||
operate and maintain the software. Operators need adjust settings to keep the
|
||||
node running, and developers need to know what options were in use when
|
||||
something goes wrong so they can diagnose and fix bugs. The legibility of a
|
||||
config file as a _human_ artifact is also thus important.
|
||||
|
||||
For example, Tendermint v0.35 moved settings related to validator private
|
||||
keys from the top-level section of the configuration file to their own
|
||||
designated `[priv-validator]` section. Although this change did not make any
|
||||
difference to the meaning of those settings, it made the organization of the
|
||||
file easier to understand, and allowed the names of the individual settings
|
||||
to be simplified (e.g., `priv-validator-key-file` became simply `key-file` in
|
||||
the new section).
|
||||
|
||||
Although such changes are "gratuitous" with respect to the software, there is
|
||||
often value in making things more legible for the humans. While there is no
|
||||
simple rule to define the line, the Potter Stewart principle can be used with
|
||||
due care.
|
||||
|
||||
Keeping these examples in mind, we can and should take reasonable steps to
|
||||
avoid churn in the configuration file across versions where we can. However, we
|
||||
must also accept that part of the reason for _having_ a config file is to allow
|
||||
us flexibility elsewhere in the design. On that basis, we should not attempt
|
||||
to be too dogmatic about config changes either. Unlike changes in the block
|
||||
protocol, for example, which affect every user of every network that adopts
|
||||
them, config changes are relatively self-contained.
|
||||
|
||||
There are few guiding principles I think we can use to strike a sensible
|
||||
balance:
|
||||
|
||||
1. **No gratuitous changes.** Aesthetic changes that do not enhance legibility,
|
||||
avert confusion, or clarity documentation, should be entirely avoided.
|
||||
|
||||
2. **Prefer mechanical changes.** Whenever it is practical, change settings in
|
||||
a way that can be updated by a tool without operator judgement. This implies
|
||||
finding safe, universal defaults for new settings, and not changing the
|
||||
default values of existing settings.
|
||||
|
||||
Even if that means we have to make multiple changes (e.g., add a new setting
|
||||
in the current version, deprecate the old one, and remove the old one in the
|
||||
next version) it's preferable if we can mechanize each step.
|
||||
|
||||
3. **Clearly signal intent.** When adding temporary or experimental settings,
|
||||
they should be clearly named and documented as such. Use long names and
|
||||
suggestive prefixes (e.g., `experimental-*`) so that they stand out when
|
||||
read in the config file or printed in logs.
|
||||
|
||||
Relatedly, using temporary or experimental settings should cause the
|
||||
software to emit diagnostic logs at runtime. These log messages should be
|
||||
easy to grep for, and should contain pointers to more complete documentation
|
||||
(say, issue numbers or URLs) that the operator can read, as well as a hint
|
||||
about when the setting is expected to become invalid. For example:
|
||||
|
||||
```
|
||||
WARNING: Websocket RPC access is deprecated and will be removed in
|
||||
Tendermint v0.37. See https://tinyurl.com/adr075 for more information.
|
||||
```
|
||||
|
||||
4. **Consider both directions.** When adding a configuration setting, take some
|
||||
time during the implementation process to think about how the setting could
|
||||
be removed, as well as how it will be rolled out. This applies even for
|
||||
settings we imagine should be permanent. Experience may cause is to rethink
|
||||
our original design intent more broadly than we expected.
|
||||
|
||||
This does not mean we have to spend a long time picking nits over the design
|
||||
of every setting; merely that we should convince ourselves we _could_ undo
|
||||
it without making too big a mess later. Even a little extra effort up front
|
||||
can sometimes save a lot.
|
||||
|
||||
## References
|
||||
|
||||
- [Tendermint `config` package][config-pkg]
|
||||
- [`confix` command-line tool][confix]
|
||||
- [`condiff` command-line tool][condiff]
|
||||
- [Configuration update plan][plan]
|
||||
- [ADR 075: RPC Event Subscription Interface][adr075]
|
||||
|
||||
[config-pkg]: https://godoc.org/github.com/tendermint/tendermint/config
|
||||
[confix]: https://github.com/tendermint/tendermint/blob/master/scripts/confix
|
||||
[condiff]: https://github.com/tendermint/tendermint/blob/master/scripts/confix/condiff
|
||||
[plan]: https://github.com/tendermint/tendermint/blob/master/scripts/confix/plan.go
|
||||
[testdata]: https://github.com/tendermint/tendermint/blob/master/scripts/confix/testdata
|
||||
[adr075]: https://github.com/tendermint/tendermint/blob/master/docs/architecture/adr-075-rpc-subscription.md
|
||||
|
||||
## Appendix: Research Notes
|
||||
|
||||
Discovering when various configuration settings were added, updated, and
|
||||
removed turns out to be surprisingly tedious. To solve this puzzle, we had to
|
||||
answer the following questions:
|
||||
|
||||
1. What changes were made between v0.x and v0.y? This is further complicated by
|
||||
cases where we have backported config changes into the middle of an earlier
|
||||
release cycle (e.g., `psql-conn` from v0.35.x into v0.34.13).
|
||||
|
||||
2. When during the development cycle were those changes made? This allows us to
|
||||
recognize features that were backported into a previous release.
|
||||
|
||||
3. What were the default values of the changed settings, and did they change at
|
||||
all during or across the release boundary?
|
||||
|
||||
Each step of the [configuration update plan][plan] is commented with a link to
|
||||
one or more PRs where that change was made. The sections below discuss how we
|
||||
found these references.
|
||||
|
||||
### Tracking Changes Across Releases
|
||||
|
||||
To figure out what changed between two releases, we built a tool called
|
||||
[`condiff`][condiff], which performs a "keyspace" diff of two TOML documents.
|
||||
This diff respects the structure of the TOML file, but ignores comments, blank
|
||||
lines, and configuration values, so that we can see what was added and removed.
|
||||
|
||||
To use it, run:
|
||||
|
||||
```shell
|
||||
go run ./scripts/confix/condiff old.toml new.toml
|
||||
```
|
||||
|
||||
This tool works on any TOML documents, but for our purposes we needed
|
||||
Tendermint `config.toml` files. The easiest way to get these is to build the
|
||||
node binary for your version of interest, run `tendermint init` on a clean home
|
||||
directory, and copy the generated config file out. The [`testdata`][testdata]
|
||||
directory for the `confix` tool has configs generated from the heads of each
|
||||
release branch from v0.31 through v0.35.
|
||||
|
||||
If you want to reproduce this yourself, it looks something like this:
|
||||
|
||||
```shell
|
||||
# Example for Tendermint v0.32.
|
||||
git checkout --track origin/v0.32.x
|
||||
go get golang.org/x/sys/unix
|
||||
go mod tidy
|
||||
make build
|
||||
rm -fr -- tmhome
|
||||
./build/tendermint --home=tmhome init
|
||||
cp tmhome/config/config.toml config-v32.toml
|
||||
```
|
||||
|
||||
Be advised that the further back you go, the more idiosyncrasies you will
|
||||
encounter. For example, Tendermint v0.31 and earlier predate Go modules (v0.31
|
||||
used dep), and lack backport branches. And you may need to do some editing of
|
||||
Makefile rules once you get back into the 20s.
|
||||
|
||||
Note that when diffing config files across the v0.34/v0.35 gap, the swap from
|
||||
`snake_case` to `kebab-case` makes it look like everything changed. The
|
||||
`condiff` tool has a `-desnake` flag that normalizes all the keys to kebab case
|
||||
in both inputs before comparison.
|
||||
|
||||
### Locating Additions and Deletions
|
||||
|
||||
To figure out when a configuration setting was added or removed, your tool of
|
||||
choice is `git bisect`. The only tricky part is finding the endpoints for the
|
||||
search. If the transition happened within a release, you can use that
|
||||
release's backport branch as the endpoint (if it has one, e.g., `v0.35.x`).
|
||||
|
||||
However, the start point can be more problematic. The backport branches are not
|
||||
ancestors of `master` or of each other, which means you need to find some point
|
||||
in history _prior_ to the change but still attached to the mainline. For recent
|
||||
releases there is a dev root (e.g., `v0.35.0-dev`, `v0.34.0-dev1`, etc.). These
|
||||
are not named consistently, but you can usually grep the output of `git tag` to
|
||||
find them.
|
||||
|
||||
In the worst case you could try starting from the root commit of the repo, but
|
||||
that turns out not to work in all cases. We've done some branching shenanigans
|
||||
over the years that mean the root is not a direct ancestor of all our release
|
||||
branches. When you find this you will probably swear a lot. I did.
|
||||
|
||||
Once you have a start and end point (say, `v0.35.0-dev` and `master`), you can
|
||||
bisect in the usual way. I use `git grep` on the `config` directory to check
|
||||
whether the case I am looking for is present. For example, to find when the
|
||||
`[fastsync]` section was removed:
|
||||
|
||||
```shell
|
||||
# Setup:
|
||||
git checkout master
|
||||
git bisect start
|
||||
git bisect bad # it's not present on tip of master.
|
||||
git bisect good v0.34.0-dev1 # it was present at the start of v0.34.
|
||||
```
|
||||
|
||||
```shell
|
||||
# Now repeat this until it gives you a specific commit:
|
||||
if git grep -q '\[fastsync\]' config ; then git bisect good ; else git bisect bad ; fi
|
||||
```
|
||||
|
||||
The above example finds where a config was removed: To find where a setting was
|
||||
added, do the same thing except reverse the sense of the test (`if ! git grep -q
|
||||
...`).
|
||||
@@ -0,0 +1,240 @@
|
||||
=======================================
|
||||
RFC 020: Tendermint Onboarding Projects
|
||||
=======================================
|
||||
|
||||
.. contents::
|
||||
:backlinks: none
|
||||
|
||||
Changelog
|
||||
---------
|
||||
|
||||
- 2022-03-30: Initial draft. (@tychoish)
|
||||
- 2022-04-25: Imported document to tendermint repository. (@tychoish)
|
||||
|
||||
Overview
|
||||
--------
|
||||
|
||||
This document describes a collection of projects that might be good for new
|
||||
engineers joining the Tendermint Core team. These projects mostly describe
|
||||
features that we'd be very excited to see land in the code base, but that are
|
||||
intentionally outside of the critical path of a release on the roadmap, and
|
||||
have the following properties that we think make good on-boarding projects:
|
||||
|
||||
- require relatively little context for the project or its history beyond a
|
||||
more isolated area of the code.
|
||||
|
||||
- provide exposure to different areas of the codebase, so new team members
|
||||
will have reason to explore the code base, build relationships with people
|
||||
on the team, and gain experience with more than one area of the system.
|
||||
|
||||
- be of moderate size, striking a healthy balance between trivial or
|
||||
mechanical changes (which provide little insight) and large intractable
|
||||
changes that require deeper insight than is available during onboarding to
|
||||
address well. A good size project should have natural touchpoints or
|
||||
check-ins.
|
||||
|
||||
Projects
|
||||
--------
|
||||
|
||||
Before diving into one of these projects, have a conversation about the
|
||||
project or aspects of Tendermint that you're excited to work on with your
|
||||
onboarding buddy. This will help make sure that these issues are still
|
||||
relevant, help you get any context, underatnding known pitfalls, and to
|
||||
confirm a high level approach or design (if relevant.) On-boarding buddies
|
||||
should be prepared to do some design work before someone joins the team.
|
||||
|
||||
The descriptions that follow provide some basic background and attempt to
|
||||
describe the user stories and the potential impact of these project.
|
||||
|
||||
E2E Test Systems
|
||||
~~~~~~~~~~~~~~~~
|
||||
|
||||
Tendermint's E2E framework makes it possible to run small test networks with
|
||||
different Tendermint configurations, and make sure that the system works. The
|
||||
tests run Tendermint in a separate binary, and the system provides some very
|
||||
high level protection against making changes that could break Tendermint in
|
||||
otherwise difficult to detect ways.
|
||||
|
||||
Working on the E2E system is a good place to get introduced to the Tendermint
|
||||
codebase, particularly for developers who are newer to Go, as the E2E
|
||||
system (generator, runner, etc.) is distinct from the rest of Tendermint and
|
||||
comparatively quite small, so it may be easier to begin making changes in this
|
||||
area. At the same time, because the E2E system exercises *all* of Tendermint,
|
||||
work in this area is a good way to get introduced to various components of the
|
||||
system.
|
||||
|
||||
Configurable E2E Workloads
|
||||
++++++++++++++++++++++++++
|
||||
|
||||
All E2E tests use the same workload (e.g. generated transactions, submitted to
|
||||
different nodes in the network,) which has been tuned empirically to provide a
|
||||
gentle but consistent parallel load that all E2E tests can pass. Ideally, the
|
||||
workload generator could be configurable to have different shapes of work
|
||||
(bursty, different transaction sizes, weighted to different nodes, etc.) and
|
||||
even perhaps further parameterized within a basic shape, which would make it
|
||||
possible to use our existing test infrastructure to answer different questions
|
||||
about the performance or capability of the system.
|
||||
|
||||
The work would involve adding a new parameter to the E2E test manifest, and
|
||||
creating an option (e.g. "legacy") for the current load generation model,
|
||||
extract configurations options for the current load generation, and then
|
||||
prototype implementations of alternate load generation, and also run some
|
||||
preliminary using the tools.
|
||||
|
||||
Byzantine E2E Workloads
|
||||
+++++++++++++++++++++++
|
||||
|
||||
There are two main kinds of integration tests in Tendermint: the E2E test
|
||||
framework, and then a collection of integration tests that masquerade as
|
||||
unit-tests. While some of this expansion of test scope is (potentially)
|
||||
inevitable, the masquerading unit tests (e.g ``consensus.byzantine_test.go``)
|
||||
end up being difficult to understand, difficult to maintain, and unreliable.
|
||||
|
||||
One solution to this, would be to modify the E2E ABCI application to allow it
|
||||
to inject byzantine behavior, and then have this be a configurable aspect of
|
||||
a test network to be able to provoke Byzantine behavior in a "real" system and
|
||||
then observe that evidence is constructed. This would make it possible to
|
||||
remove the legacy tests entirely once the new tests have proven themselves.
|
||||
|
||||
Abstract Orchestration Framework
|
||||
++++++++++++++++++++++++++++++++
|
||||
|
||||
The orchestration of e2e test processes is presently done using docker
|
||||
compose, which works well, but has proven a bit limiting as all processes need
|
||||
to run on a single machine, and the log aggregation functions are confusing at
|
||||
best.
|
||||
|
||||
This project would replace the current orchestration with something more
|
||||
generic, potentially maintaining the current system, but also allowing the e2e
|
||||
tests to manage processes using k8s. There are a few "local" k8s frameworks
|
||||
(e.g. kind and k3s,) which might be able to be useful for our current testing
|
||||
model, but hopefully, we could use this new implementation with other k8s
|
||||
systems for more flexible distribute test orchestration.
|
||||
|
||||
Improve Operationalize Experience of ``run-multiple.sh``
|
||||
++++++++++++++++++++++++++++++++++++++++++++++++++++++++
|
||||
|
||||
The e2e test runner currently runs a single test, and in most cases we manage
|
||||
the test cases using a shell script that ensure cleanup of entire test
|
||||
suites. This is a bit difficult to maintain and makes reproduction of test
|
||||
cases more awkward than it should be. The e2e ``runner`` itself should provide
|
||||
equivalent functionality to ``run-multiple.sh``: ensure cleanup of test cases,
|
||||
collect and process output, and be able to manage entire suites of cases.
|
||||
|
||||
It might also be useful to implement an e2e test orchestrator that runs all
|
||||
tendermint instances in a single process, using "real" networks for faster
|
||||
feedback and iteration during development.
|
||||
|
||||
In addition to being a bit easier to maintain, having a more capable runner
|
||||
implementation would make it easier to collect data from test runs, improve
|
||||
debugability and reporting.
|
||||
|
||||
Fan-Out For CI E2E Tests
|
||||
++++++++++++++++++++++++
|
||||
|
||||
While there are some parallelism in the execution of e2e tests, each e2e test
|
||||
job must build a tendermint e2e image, which takes about 5 minutes of CPU time
|
||||
per-task, which given the size of each of the runs.
|
||||
|
||||
We'd like to be able to reduce the amount of overhead per-e2e tests while
|
||||
keeping the cycle time for working with the tests very low, while also
|
||||
maintaining a reasonable level of test coverage. This is an impossible
|
||||
tradeoff, in some ways, and the percentage of overhead at the moment is large
|
||||
enough that we can make some material progress with a moderate amount of time.
|
||||
|
||||
Most of this work has to do with modifying github actions configuration and
|
||||
e2e artifact (docker) building to reduce redundant work. Eventually, when we
|
||||
can drop the requirement for CGo storage engines, it will be possible to move
|
||||
(cross) compile tendermint locally, and then inject the binary into the docker
|
||||
container, which would reduce a lot of the build-time complexity, although we
|
||||
can move more in this direction or have runtime flags to disable CGo
|
||||
dependencies for local development.
|
||||
|
||||
Remove Panics
|
||||
~~~~~~~~~~~~~
|
||||
|
||||
There are lots of places in the code base which can panic, and would not be
|
||||
particularly well handled. While in some cases, panics are the right answer,
|
||||
in many cases the panics were just added to simplify downstream error
|
||||
checking, and could easily be converted to errors.
|
||||
|
||||
The `Don't Panic RFC
|
||||
<https://github.com/tendermint/tendermint/blob/master/docs/rfc/rfc-008-do-not-panic.MD>`_
|
||||
covers some of the background and approach.
|
||||
|
||||
While the changes are in this project are relatively rote, this will provide
|
||||
exposure to lots of different areas of the codebase as well as insight into
|
||||
how different areas of the codebase interact with eachother, as well as
|
||||
experience with the test suites and infrastructure.
|
||||
|
||||
Implement more Expressive ABCI Applications
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Tendermint maintains two very simple ABCI applications (a KV application used
|
||||
for basic testing, and slightly more advanced test application used in the
|
||||
end-to-end tests). Writing an application would provide a new engineer with
|
||||
useful experiences using Tendermint that mirrors the expierence of downstream
|
||||
users.
|
||||
|
||||
This is more of an exploratory project, but could include providing common
|
||||
interfaces on top of Tendermint consensus for other well known protocols or
|
||||
tools (e.g. ``etcd``) or a DNS server or some other tool.
|
||||
|
||||
Self-Regulating Reactors
|
||||
~~~~~~~~~~~~~~~~~~~~~~~~
|
||||
|
||||
Currently reactors (the internal processes that are responsible for the higher
|
||||
level behavior of Tendermint) can be started and stopped, but have no
|
||||
provision for being paused. These additional semantics may allow Tendermint to
|
||||
pause reactors (and avoid processing their messhages, etc.) and allow better
|
||||
coordination in the future.
|
||||
|
||||
While this is a big project, it's possible to break this apart into many
|
||||
smaller projects: make p2p channels pauseable, add pause/UN-pause hooks to the
|
||||
service implementation and machinery, and finally to modify the reactor
|
||||
implementations to take advantage of these additional semantics
|
||||
|
||||
This project would give an engineer some exposure to the p2p layer of the
|
||||
code, as well as to various aspects of the reactor implementations.
|
||||
|
||||
Metrics
|
||||
~~~~~~~
|
||||
|
||||
Tendermint has a metrics system that is relatively underutilized, and figuring
|
||||
out ways to capture and organize the metrics to provide value to users might
|
||||
provide an interesting set of projects for new engineers on Tendermint.
|
||||
|
||||
Convert Logs to Metrics
|
||||
+++++++++++++++++++++++
|
||||
|
||||
Because the tendermint logs tend to be quite verbose and not particularly
|
||||
actionable, most users largely ignore the logging or run at very low
|
||||
verbosity. While the log statements in the code do describe useful events,
|
||||
taken as a whole the system is not particularly tractable, and particularly at
|
||||
the Debug level, not useful. One solution to this problem is to identify log
|
||||
messages that might be (e.g. increment a counter for certian kinds of errors)
|
||||
|
||||
One approach might be to look at various logging statements, particularly
|
||||
debug statements or errors that are logged but not returned, and see if
|
||||
they're convertable to counters or other metrics.
|
||||
|
||||
Expose Metrics to Tests
|
||||
+++++++++++++++++++++++
|
||||
|
||||
The existing Tendermint test suites replace the metrics infrastructure with
|
||||
no-op implementations, which means that tests can neither verify that metrics
|
||||
are ever recorded, nor can tests use metrics to observe events in the
|
||||
system. Writing an implementation, for testing, that makes it possible to
|
||||
record metrics and provides an API for introspecting this data, as well as
|
||||
potentially writing tests that take advantage of this type, could be useful.
|
||||
|
||||
Logging Metrics
|
||||
+++++++++++++++
|
||||
|
||||
In some systems, the logging system itself can provide some interesting
|
||||
insights for operators: having metrics that track the number of messages at
|
||||
different levels as well as the total number of messages, can act as a canary
|
||||
for the system as a whole.
|
||||
|
||||
This should be achievable by adding an interceptor layer within the logging
|
||||
package itself that can add metrics to the existing system.
|
||||
@@ -0,0 +1,266 @@
|
||||
# RFC 021: The Future of the Socket Protocol
|
||||
|
||||
## Changelog
|
||||
|
||||
- 19-May-2022: Initial draft (@creachadair)
|
||||
- 19-Jul-2022: Converted from ADR to RFC (@creachadair)
|
||||
|
||||
## Abstract
|
||||
|
||||
This RFC captures some technical discussion about the ABCI socket protocol that
|
||||
was originally documented to solicit an architectural decision. This topic was
|
||||
not high-enough priority as of this writing to justify making a final decision.
|
||||
|
||||
For that reason, the text of this RFC has the general structure of an ADR, but
|
||||
should be viewed primarily as a record of the issue for future reference.
|
||||
|
||||
## Background
|
||||
|
||||
The [Application Blockchain Interface (ABCI)][abci] is a client-server protocol
|
||||
used by the Tendermint consensus engine to communicate with the application on
|
||||
whose behalf it performs state replication. There are currently three transport
|
||||
options available for ABCI applications:
|
||||
|
||||
1. **In-process**: Applications written in Go can be linked directly into the
|
||||
same binary as the consensus node. Such applications use a "local" ABCI
|
||||
connection, which exposes application methods to the node as direct function
|
||||
calls.
|
||||
|
||||
2. **Socket protocol**: Out-of-process applications may export the ABCI service
|
||||
via a custom socket protocol that sends requests and responses over a
|
||||
Unix-domain or TCP socket connection as length-prefixed protocol buffers.
|
||||
In Tendermint, this is handled by the [socket client][socket-client].
|
||||
|
||||
3. **gRPC**: Out-of-process applications may export the ABCI service via gRPC.
|
||||
In Tendermint, this is handled by the [gRPC client][grpc-client].
|
||||
|
||||
Both the out-of-process options (2) and (3) have a long history in Tendermint.
|
||||
The beginnings of the gRPC client were added in [May 2016][abci-start] when
|
||||
ABCI was still hosted in a separate repository, and the socket client (formerly
|
||||
called the "remote client") was part of ABCI from its inception in November
|
||||
2015.
|
||||
|
||||
At that time when ABCI was first being developed, the gRPC project was very new
|
||||
(it launched Q4 2015) and it was not an obvious choice for use in Tendermint.
|
||||
It took a while before the language coverage and quality of gRPC reached a
|
||||
point where it could be a viable solution for out-of-process applications. For
|
||||
that reason, it made sense for the initial design of ABCI to focus on a custom
|
||||
protocol for out-of-process applications.
|
||||
|
||||
## Problem Statement
|
||||
|
||||
For practical reasons, ABCI needs an interprocess communication option to
|
||||
support applications not written in Go. The two practical options are RPC and
|
||||
FFI, and for operational reasons an RPC mechanism makes more sense.
|
||||
|
||||
The socket protocol has not changed all that substantially since its original
|
||||
design, and has the advantage of being simple to implement in almost any
|
||||
reasonable language. However, its simplicity includes some limitations that
|
||||
have had a negative impact on the stability and performance of out-of-process
|
||||
applications using it. In particular:
|
||||
|
||||
- The protocol lacks request identifiers, so the client and server must return
|
||||
responses in strict FIFO order. Even if the client issues requests that have
|
||||
no dependency on each other, the protocol has no way except order of issue to
|
||||
map responses to requests.
|
||||
|
||||
This reduces (in some cases substantially) the concurrency an application can
|
||||
exploit, since the parallelism of requests in flight is gated by the slowest
|
||||
active request at any moment. There have been complaints from some network
|
||||
operators on that basis.
|
||||
|
||||
- The protocol lacks method identifiers, so the only way for the client and
|
||||
server to understand which operation is requested is to dispatch on the type
|
||||
of the request and response payloads. For responses, this means that [any
|
||||
error condition is terminal not only to the request, but to the entire ABCI
|
||||
client](https://github.com/tendermint/tendermint/blob/master/abci/client/socket_client.go#L149).
|
||||
|
||||
The historical intent of terminating for any error seems to have been that
|
||||
all ABCI errors are unrecoverable and hence protocol fatal <!-- markdown-link-check-disable-next-line -->
|
||||
(see [Note 1](#note1)). In practice, however, this greatly complicates
|
||||
debugging a faulty node, since the only way to respond to errors is to panic
|
||||
the node which loses valuable context that could have been logged.
|
||||
|
||||
- There are subtle concurrency management dependencies between the client and
|
||||
the server that are not clearly documented anywhere, and it is very easy for
|
||||
small changes in both the client and the server to lead to tricky deadlocks,
|
||||
panics, race conditions, and slowdowns. As a recent example of this, see
|
||||
https://github.com/tendermint/tendermint/pull/8581.
|
||||
|
||||
These limitations are fixable, but one important question is whether it is
|
||||
worthwhile to fix them. We can add request and method identifiers, for
|
||||
example, but doing so would be a breaking change to the protocol requiring
|
||||
every application using it to update. If applications have to migrate anyway,
|
||||
the stability and language coverage of gRPC have improved a lot, and today it
|
||||
is probably simpler to set up and maintain an application using gRPC transport
|
||||
than to reimplement the Tendermint socket protocol.
|
||||
|
||||
Moreover, gRPC addresses all the above issues out-of-the-box, and requires
|
||||
(much) less custom code for both the server (i.e., the application) and the
|
||||
client. The project is well-funded and widely-used, which makes it a safe bet
|
||||
for a dependency.
|
||||
|
||||
## Decision
|
||||
|
||||
There is a set of related alternatives to consider:
|
||||
|
||||
- Question 1: Designate a single IPC standard for out-of-process applications?
|
||||
|
||||
Claim: We should converge on one (and only one) IPC option for out-of-process
|
||||
applications. We should choose an option that, after a suitable period of
|
||||
deprecation for alternatives, will address most or all the highest-impact
|
||||
uses of Tendermint. Maintaining multiple options increases the surface area
|
||||
for bugs and vulnerabilities, and we should not have multiple options for
|
||||
basic interfaces without a clear and well-documented reason.
|
||||
|
||||
- Question 2a: Choose gRPC and deprecate/remove the socket protocol?
|
||||
|
||||
Claim: Maintaining and improving a custom RPC protocol is a substantial
|
||||
project and not directly relevant to the requirements of consensus. We would
|
||||
be better served by depending on a well-maintained open-source library like
|
||||
gRPC.
|
||||
|
||||
- Question 2b: Improve the socket protocol and deprecate/remove gRPC?
|
||||
|
||||
Claim: If we find meaningful advantages to maintaining our own custom RPC
|
||||
protocol in Tendermint, we should treat it as a first-class project within
|
||||
the core and invest in making it good enough that we do not require other
|
||||
options.
|
||||
|
||||
**One important consideration** when discussing these questions is that _any
|
||||
outcome which includes keeping the socket protocol will have eventual migration
|
||||
impacts for out-of-process applications_ regardless. To fix the limitations of
|
||||
the socket protocol as it is currently designed will require making _breaking
|
||||
changes_ to the protocol. So, while we may put off a migration cost for
|
||||
out-of-process applications by retaining the socket protocol in the short term,
|
||||
we will eventually have to pay those costs to fix the problems in its current
|
||||
design.
|
||||
|
||||
## Detailed Design
|
||||
|
||||
1. If we choose to standardize on gRPC, the main work in Tendermint core will
|
||||
be removing and cleaning up the code for the socket client and server.
|
||||
|
||||
Besides the code cleanup, we will also need to clearly document a
|
||||
deprecation schedule, and invest time in making the migration easier for
|
||||
applications currently using the socket protocol.
|
||||
|
||||
> **Point for discussion:** Migrating from the socket protocol to gRPC
|
||||
> should mostly be a plumbing change, as long as we do it during a release
|
||||
> in which we are not making other breaking changes to ABCI. However, the
|
||||
> effort may be more or less depending on how gRPC integration works in the
|
||||
> application's implementation language, and would have to be sure networks
|
||||
> have plenty of time not only to make the change but to verify that it
|
||||
> preserves the function of the network.
|
||||
>
|
||||
> What questions should we be asking node operators and application
|
||||
> developers to understand the migration costs better?
|
||||
|
||||
2. If we choose to keep only the socket protocol, we will need to follow up
|
||||
with a more detailed design for extending and upgrading the protocol to fix
|
||||
the existing performance and operational issues with the protocol.
|
||||
|
||||
Moreover, since the gRPC interface has been around for a long time we will
|
||||
also need a deprecation plan for it.
|
||||
|
||||
3. If we choose to keep both options, we will still need to do all the work of
|
||||
(2), but the gRPC implementation should not require any immediate changes.
|
||||
|
||||
|
||||
## Alternatives Considered
|
||||
|
||||
- **FFI**. Another approach we could take is to use a C-based FFI interface so
|
||||
that applications written in other languages are linked directly with the
|
||||
consensus node, an option currently only available for Go applications.
|
||||
|
||||
An FFI interface is possible for a lot of languages, but FFI support varies
|
||||
widely in coverage and quality across languages and the points of friction
|
||||
can be tricky to work around. Moreover, it's much harder to add FFI support
|
||||
to a language where it's missing after-the-fact for an application developer.
|
||||
|
||||
Although a basic FFI interface is not too difficult on the Go side, the C
|
||||
shims for an FFI can get complicated if there's a lot of variability in the
|
||||
runtime environment on the other end.
|
||||
|
||||
If we want to have one answer for non-Go applications, we are better off
|
||||
picking an IPC-based solution (whether that's gRPC or an extension of our
|
||||
custom socket protocol or something else).
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Standardize on gRPC**
|
||||
|
||||
- ✅ Addresses existing performance and operational issues.
|
||||
- ✅ Replaces custom code with a well-maintained widely-used library.
|
||||
- ✅ Aligns with Cosmos SDK, which already uses gRPC extensively.
|
||||
- ✅ Aligns with priv validator interface, for which the socket protocol is already deprecated for gRPC.
|
||||
- ❓ Applications will be hard to implement in a language without gRPC support.
|
||||
- ⛔ All users of the socket protocol have to migrate to gRPC, and we believe most current out-of-process applications use the socket protocol.
|
||||
|
||||
- **Standardize on socket protocol**
|
||||
|
||||
- ✅ Less immediate impact for existing users (but see below).
|
||||
- ✅ Simplifies ABCI API surface by removing gRPC.
|
||||
- ❓ Users of the socket protocol will have a (smaller) migration.
|
||||
- ❓ Potentially easier to implement for languages that do not have support.
|
||||
- ⛔ Need to do all the work to fix the socket protocol (which will require existing users to update anyway later).
|
||||
- ⛔ Ongoing maintenance burden for per-language server implementations.
|
||||
|
||||
- **Keep both options**
|
||||
|
||||
- ✅ Less immediate impact for existing users (but see below).
|
||||
- ❓ Users of the socket protocol will have a (smaller) migration.
|
||||
- ⛔ Still need to do all the work to fix the socket protocol (which will require existing users to update anyway later).
|
||||
- ⛔ Requires ongoing maintenance and support of both gRPC and socket protocol integrations.
|
||||
|
||||
|
||||
## References
|
||||
|
||||
- [Application Blockchain Interface (ABCI)][abci]
|
||||
- [Tendermint ABCI socket client][socket-client]
|
||||
- [Tendermint ABCI gRPC client][grpc-client]
|
||||
- [Initial commit of gRPC client][abci-start]
|
||||
|
||||
[abci]: https://github.com/tendermint/spec/tree/master/spec/abci
|
||||
[socket-client]: https://github.com/tendermint/tendermint/blob/master/abci/client/socket_client.go
|
||||
[socket-server]: https://github.com/tendermint/tendermint/blob/master/abci/server/socket_server.go
|
||||
[grpc-client]: https://github.com/tendermint/tendermint/blob/master/abci/client/grpc_client.go
|
||||
[abci-start]: https://github.com/tendermint/abci/commit/1ab3c747182aaa38418258679c667090c2bb1e0d
|
||||
|
||||
## Notes
|
||||
|
||||
- <a id=note1></a>**Note 1**: The choice to make all ABCI errors protocol-fatal
|
||||
was intended to avoid the risk that recovering an application error could
|
||||
cause application state to diverge. Divergence can break consensus, so it's
|
||||
essential to avoid it.
|
||||
|
||||
This is a sound principle, but conflates protocol errors with "mechanical"
|
||||
errors such as timeouts, resoures exhaustion, failed connections, and so on.
|
||||
Because the protocol has no way to distinguish these conditions, the only way
|
||||
for an application to report an error is to panic or crash.
|
||||
|
||||
Whether a node is running in the same process as the application or as a
|
||||
separate process, application errors should not be suppressed or hidden.
|
||||
However, it's important to ensure that errors are handled at a consistent and
|
||||
well-defined point in the protocol: Having the application panic or crash
|
||||
rather than reporting an error means the node sees different results
|
||||
depending on whether the application runs in-process or out-of-process, even
|
||||
if the application logic is otherwise identical.
|
||||
|
||||
## Appendix: Known Implementations of ABCI Socket Protocol
|
||||
|
||||
This is a list of known implementations of the Tendermint custom socket
|
||||
protocol. Note that in most cases I have not checked how complete or correct
|
||||
these implementations are; these are based on search results and a cursory
|
||||
visual inspection.
|
||||
|
||||
- Tendermint Core (Go): [client][socket-client], [server][socket-server]
|
||||
- Informal Systems [tendermint-rs](https://github.com/informalsystems/tendermint-rs) (Rust): [client](https://github.com/informalsystems/tendermint-rs/blob/master/abci/src/client.rs), [server](https://github.com/informalsystems/tendermint-rs/blob/master/abci/src/server.rs)
|
||||
- Tendermint [js-abci](https://github.com/tendermint/js-abci) (JS): [server](https://github.com/tendermint/js-abci/blob/master/src/server.js)
|
||||
- [Hotmoka](https://github.com/Hotmoka/hotmoka) ABCI (Java): [server](https://github.com/Hotmoka/hotmoka/blob/master/io-hotmoka-tendermint-abci/src/main/java/io/hotmoka/tendermint_abci/Server.java)
|
||||
- [Tower ABCI](https://github.com/penumbra-zone/tower-abci) (Rust): [server](https://github.com/penumbra-zone/tower-abci/blob/main/src/server.rs)
|
||||
- [abci-host](https://github.com/datopia/abci-host) (Clojure): [server](https://github.com/datopia/abci-host/blob/master/src/abci/host.clj)
|
||||
- [abci_server](https://github.com/KrzysiekJ/abci_server) (Erlang): [server](https://github.com/KrzysiekJ/abci_server/blob/master/src/abci_server.erl)
|
||||
- [py-abci](https://github.com/davebryson/py-abci) (Python): [server](https://github.com/davebryson/py-abci/blob/master/src/abci/server.py)
|
||||
- [scala-tendermint-server](https://github.com/intechsa/scala-tendermint-server) (Scala): [server](https://github.com/InTechSA/scala-tendermint-server/blob/master/src/main/scala/lu/intech/tendermint/Server.scala)
|
||||
- [kepler](https://github.com/f-o-a-m/kepler) (Rust): [server](https://github.com/f-o-a-m/kepler/blob/master/hs-abci-server/src/Network/ABCI/Server.hs)
|
||||
@@ -0,0 +1,35 @@
|
||||
# RFC {RFC-NUMBER}: {TITLE}
|
||||
|
||||
## Changelog
|
||||
|
||||
- {date}: {changelog}
|
||||
|
||||
## Abstract
|
||||
|
||||
> A brief high-level synopsis of the topic of discussion for this RFC, ideally
|
||||
> just a few sentences. This should help the reader quickly decide whether the
|
||||
> rest of the discussion is relevant to their interest.
|
||||
|
||||
## Background
|
||||
|
||||
> Any context or orientation needed for a reader to understand and participate
|
||||
> in the substance of the Discussion. If necessary, this section may include
|
||||
> links to other documentation or sources rather than restating existing
|
||||
> material, but should provide enough detail that the reader can tell what they
|
||||
> need to read to be up-to-date.
|
||||
|
||||
### References
|
||||
|
||||
> Links to external materials needed to follow the discussion may be added here.
|
||||
>
|
||||
> In addition, if the discussion in a request for comments leads to any design
|
||||
> decisions, it may be helpful to add links to the ADR documents here after the
|
||||
> discussion has settled.
|
||||
|
||||
## Discussion
|
||||
|
||||
> This section contains the core of the discussion.
|
||||
>
|
||||
> There is no fixed format for this section, but ideally changes to this
|
||||
> section should be updated before merging to reflect any discussion that took
|
||||
> place on the PR that made those changes.
|
||||
Reference in New Issue
Block a user