mirror of
https://github.com/tendermint/tendermint.git
synced 2026-07-20 15:02:33 +00:00
Documentation of p2p layer in Tendermint v0.34 (#9348)
* spec: overview of p2p in v0.34 (#9120) * Iniital comments on v0.34 p2p * Added conf, updated text * Moved everything to spec * Update README.md * spec: overview of the p2p implementation in v0.34 (#9126) * Spec: p2p v0.34 doc, switch initial documentation * Spec: p2p v0.34 doc, list of source files * Spec: p2p v0.34 doc, transport documentation * Spec: p2p v0.34 doc, transport error handling * Spec: p2p v0.34 doc, PEX initial documentation * PEX protocol documentation is a separated file * PEX reactor documentation with a general documentation, including the address book and its role as (outbound) peer manager. * Spec: p2p v0.34 doc, PEX protocol documentation * Spec: p2p v0.34 doc, PEX protocol on seed nodes * Spec: p2p v0.34 doc, address book * Spec: p2p v0.34 doc, address book, more details * Spec: p2p v0.34 doc, address book persistence * Spec: p2p v0.34 doc, address book random samples * Spec: p2p v0.34 doc, status of this documentation * Spec: p2p v0.34 doc, pex reactor documentation * Spec: p2p v0.34 doc, addressing PR #9126 comments Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> * Spec: p2p v0.34 doc, peer manager, outbound peers Co-authored-by: Daniel Cason <cason@gandria> Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> * spec:p2p v0.34 introduction (#9319) * restructure README.md initial * Fix typos * Reorganization * spec: overview of p2p in v0.34 (#9120) * Iniital comments on v0.34 p2p * Added conf, updated text * Moved everything to spec * Update README.md * spec: overview of the p2p implementation in v0.34 (#9126) * Spec: p2p v0.34 doc, switch initial documentation * Spec: p2p v0.34 doc, list of source files * Spec: p2p v0.34 doc, transport documentation * Spec: p2p v0.34 doc, transport error handling * Spec: p2p v0.34 doc, PEX initial documentation * PEX protocol documentation is a separated file * PEX reactor documentation with a general documentation, including the address book and its role as (outbound) peer manager. * Spec: p2p v0.34 doc, PEX protocol documentation * Spec: p2p v0.34 doc, PEX protocol on seed nodes * Spec: p2p v0.34 doc, address book * Spec: p2p v0.34 doc, address book, more details * Spec: p2p v0.34 doc, address book persistence * Spec: p2p v0.34 doc, address book random samples * Spec: p2p v0.34 doc, status of this documentation * Spec: p2p v0.34 doc, pex reactor documentation * Spec: p2p v0.34 doc, addressing PR #9126 comments Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> * Spec: p2p v0.34 doc, peer manager, outbound peers Co-authored-by: Daniel Cason <cason@gandria> Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> * spec:p2p v0.34 introduction (#9319) * restructure README.md initial * Fix typos * Reorganization * spec: p2p v0.34, addressbook review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * spec: p2p v0.34, peer manager review * Filled config description * spec: p2p v0.34, transport review * spec: p2p v0.34, switch review * spec: p2p v0.34, overview, first version * spec: p2p v0.34, peer manager review * spec: p2p v0.34, shorter readme * Configuration update * Configuration update * Shortened README * spec: p2p v0.34, readme intro * spec: p2p v0.34, readme contents * spec: p2p v0.34, readme references * spec: p2p readme points to v0.34 * spec: p2p, v0.34, fixing brokend markdown links * Makrdown fix * Apply suggestions from code review Co-authored-by: Adi Seredinschi <adizere@gmail.com> Co-authored-by: Zarko Milosevic <zarko@informal.systems> * spec: p2p v0.34, address book new intro * spec: p2p v0.34, address book buckets summary * spec: p2p v0.34, peer manager, issue link * spec: p2p v0.34, fixing links * spec: p2p v0.34, addressing comments from reviews * spec: p2p v0.34, addressing comments from reviews * Apply suggestions from Jasmina's code review Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> * spec: p2p v0.34, addressing comments from reviews * Apply suggestions from code review Co-authored-by: Sergio Mena <sergio@informal.systems> * Apply suggestions from code review Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> Co-authored-by: Sergio Mena <sergio@informal.systems> * spec: p2p, v0.34, address book section reorganized * spec: p2p, v0.34, addressing review comments * Typos * Typo * spec: p2p, v0.34, address book markbad Co-authored-by: Jasmina Malicevic <jasmina.dustinac@gmail.com> Co-authored-by: Daniel Cason <cason@gandria> Co-authored-by: Adi Seredinschi <adizere@gmail.com> Co-authored-by: Zarko Milosevic <zarko@informal.systems> Co-authored-by: Sergio Mena <sergio@informal.systems>
This commit is contained in:
@@ -4,3 +4,10 @@ parent:
|
||||
title: P2P
|
||||
order: 6
|
||||
---
|
||||
|
||||
# Peer-to-Peer Communication
|
||||
|
||||
The operation of the p2p adopted in production Tendermint networks is [HERE](./v0.34/).
|
||||
|
||||
> This is part of an ongoing [effort](https://github.com/tendermint/tendermint/issues/9089)
|
||||
> to produce a high-level specification of the operation of the p2p layer.
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
# Peer-to-Peer Communication
|
||||
|
||||
This document describes the implementation of the peer-to-peer (p2p)
|
||||
communication layer in Tendermint.
|
||||
|
||||
It is part of an [effort](https://github.com/tendermint/tendermint/issues/9089)
|
||||
to produce a high-level specification of the operation of the p2p layer adopted
|
||||
in production Tendermint networks.
|
||||
|
||||
This documentation, therefore, considers the releases `0.34.*` of Tendermint, more
|
||||
specifically, the branch [`v0.34.x`](https://github.com/tendermint/tendermint/tree/v0.34.x)
|
||||
of this repository.
|
||||
|
||||
## Overview
|
||||
|
||||
A Tendermint network is composed of multiple Tendermint instances, hereafter
|
||||
called **nodes**, that interact by exchanging messages.
|
||||
|
||||
Tendermint assumes a partially-connected network model.
|
||||
This means that a node is not assumed to be directly connected to every other
|
||||
node in the network.
|
||||
Instead, each node is directly connected to a subset of other nodes in the
|
||||
network, hereafter called its **peers**.
|
||||
|
||||
The peer-to-peer (p2p) communication layer is responsible for establishing
|
||||
connections between nodes in a Tendermint network,
|
||||
for managing the communication between a node and its peers,
|
||||
and for intermediating the exchange of messages between peers in Tendermint protocols.
|
||||
|
||||
## Contents
|
||||
|
||||
The documentation follows the organization of the `p2p` package of Tendermint,
|
||||
which implements the following abstractions:
|
||||
|
||||
- [Transport](./transport.md): establishes secure and authenticated
|
||||
connections with peers;
|
||||
- [Switch](./switch.md): responsible for dialing peers and accepting
|
||||
connections from peers, for managing established connections, and for
|
||||
routing messages between the reactors and peers,
|
||||
that is, between local and remote instances of the Tendermint protocols;
|
||||
- [PEX Reactor](./pex.md): a reactor is the implementation of a protocol which
|
||||
exchanges messages through the p2p layer. The PEX reactor manages the [Address Book](./addressbook.md) and implements both the [PEX protocol](./pex-protocol.md) and the [Peer Manager](./peer_manager.md) role.
|
||||
- [Peer Exchange protocol](./pex-protocol.md): enables nodes to exchange peer addresses, thus implementing a peer discovery service;
|
||||
- [Address Book](./addressbook.md): stores discovered peer addresses and
|
||||
quality metrics associated to peers with which the node has interacted;
|
||||
- [Peer Manager](./peer_manager.md): defines when and to which peers a node
|
||||
should dial, in order to establish outbound connections;
|
||||
- Finally, [Types](./types.md) and [Configuration](./configuration.md) provide
|
||||
a list of existing types and configuration parameters used by the p2p layer implementation.
|
||||
|
||||
## Further References
|
||||
|
||||
Existing documentation referring to the p2p layer:
|
||||
|
||||
- https://github.com/tendermint/tendermint/tree/main/spec/p2p: p2p-related
|
||||
configuration flags; overview of connections, peer instances, and reactors;
|
||||
overview of peer discovery and node types; peer identity, secure connections
|
||||
and peer authentication handshake.
|
||||
- https://github.com/tendermint/tendermint/tree/main/spec/p2p/messages: message
|
||||
types and channel IDs of Block Sync, Mempool, Evidence, State Sync, PEX, and
|
||||
Consensus reactors.
|
||||
- https://docs.tendermint.com/v0.34/tendermint-core: the p2p layer
|
||||
configuration and operation is documented in several pages.
|
||||
This content is not necessarily up-to-date, some settings and concepts may
|
||||
refer to the release `v0.35`, that was [discontinued][v35postmorten].
|
||||
- https://github.com/tendermint/tendermint/tree/master/docs/tendermint-core/pex:
|
||||
peer types, peer discovery, peer management overview, address book and peer
|
||||
ranking. This documentation refers to the release `v0.35`, that was [discontinued][v35postmorten].
|
||||
|
||||
[v35postmorten]: https://interchain-io.medium.com/discontinuing-tendermint-v0-35-a-postmortem-on-the-new-networking-layer-3696c811dabc
|
||||
@@ -0,0 +1,367 @@
|
||||
# Address Book
|
||||
|
||||
The address book tracks information about peers, i.e., about other nodes in the network.
|
||||
|
||||
The primary information stored in the address book are peer addresses.
|
||||
A peer address is composed by a node ID and a network address; a network
|
||||
address is composed by an IP address or a DNS name plus a port number.
|
||||
The same node ID can be associated to multiple network addresses.
|
||||
|
||||
There are two sources for the addresses stored in the address book.
|
||||
The [Peer Exchange protocol](./pex-protocol.md) stores in the address book
|
||||
the peer addresses it discovers, i.e., it learns from connected peers.
|
||||
And the [Switch](./switch.md) registers the addresses of peers with which it
|
||||
has interacted: to which it has dialed or from which it has accepted a
|
||||
connection.
|
||||
|
||||
The address book also records additional information about peers with which the
|
||||
node has interacted, from which is possible to rank peers.
|
||||
The Switch reports [connection attempts](#dial-attempts) to a peer address; too
|
||||
much failed attempts indicate that a peer address is invalid.
|
||||
Reactors, in they turn, report a peer as [good](#good-peers) when it behaves as
|
||||
expected, or as a [bad peer](#bad-peers), when it misbehaves.
|
||||
|
||||
There are two entities that retrieve peer addresses from the address book.
|
||||
The [Peer Manager](./peer_manager.md) retrieves peer addresses to dial, so to
|
||||
establish outbound connections.
|
||||
This selection is random, but has a configurable bias towards peers that have
|
||||
been marked as good peers.
|
||||
The [Peer Exchange protocol](./pex-protocol.md) retrieves random samples of
|
||||
addresses to offer (send) to peers.
|
||||
This selection is also random but it includes, in particular for nodes that
|
||||
operate in seed mode, some bias toward peers marked as good ones.
|
||||
|
||||
## Buckets
|
||||
|
||||
Peer addresses are stored in buckets.
|
||||
There are buckets for new addresses and buckets for old addresses.
|
||||
The buckets for new addresses store addresses of peers about which the node
|
||||
does not have much information; the first address registered for a peer ID is
|
||||
always stored in a bucket for new addresses.
|
||||
The buckets for old addresses store addresses of peers with which the node has
|
||||
interacted and that were reported as [good peers](#good-peers) by a reactor.
|
||||
An old address therefore can be seen as an alias for a good address.
|
||||
|
||||
> Note that new addresses does not mean bad addresses.
|
||||
> The addresses of peers marked as [bad peers](#bad-peers) are removed from the
|
||||
> buckets where they are stored, and temporarily kept in a table of banned peers.
|
||||
|
||||
The number of buckets is fixed and there are more buckets for new addresses
|
||||
(`256`) than buckets for old addresses (`64`), a ratio of 4:1.
|
||||
Each bucket can store up to `64` addresses.
|
||||
When a bucket becomes full, the peer address with the lowest ranking is removed
|
||||
from the bucket.
|
||||
The first choice is to remove bad addresses, with multiple failed attempts
|
||||
associated.
|
||||
In the absence of those, the *oldest* address in the bucket is removed, i.e.,
|
||||
the address with the oldest last attempt to dial.
|
||||
|
||||
When a bucket for old addresses becomes full, the lowest-ranked peer address in
|
||||
the bucket is moved to a bucket of new addresses.
|
||||
When a bucket for new addresses becomes full, the lowest-ranked peer address in
|
||||
the bucket is removed from the address book.
|
||||
In other words, exceeding old or good addresses are downgraded to new
|
||||
addresses, while exceeding new addresses are dropped.
|
||||
|
||||
The bucket that stores an `address` is defined by the following two methods,
|
||||
for new and old addresses:
|
||||
|
||||
- `calcNewBucket(address, source) = hash(key + groupKey(source) + hash(key + groupKey(address) + groupKey(source)) % newBucketsPerGroup) % newBucketCount`
|
||||
- `calcOldBucket(address) = hash(key + groupKey(address) + hash(key + address) % oldBucketsPerGroup) % oldBucketCount`
|
||||
|
||||
The `key` is a fixed random 96-bit (8-byte) string.
|
||||
The `groupKey` for an address is a string representing its network group.
|
||||
The `source` of an address is the address of the peer from which we learn the
|
||||
address..
|
||||
The first (internal) hash is reduced to an integer up to `newBucketsPerGroup =
|
||||
32`, for new addresses, and `oldBucketsPerGroup = 4`, for old addresses.
|
||||
The second (external) hash is reduced to bucket indexes, in the interval from 0
|
||||
to the number of new (`newBucketCount = 256`) or old (`oldBucketCount = 64`) buckets.
|
||||
|
||||
Notice that new addresses with sources from the same network group are more
|
||||
likely to end up in the same bucket, therefore to competing for it.
|
||||
For old address, instead, two addresses are more likely to end up in the same
|
||||
bucket when they belong to the same network group.
|
||||
|
||||
## Adding addresses
|
||||
|
||||
The `AddAddress` method adds the address of a peer to the address book.
|
||||
|
||||
The added address is associated to a *source* address, which identifies the
|
||||
node from which the peer address was learned.
|
||||
|
||||
Addresses are added to the address book in the following situations:
|
||||
|
||||
1. When a peer address is learned via PEX protocol, having the sender
|
||||
of the PEX message as its source
|
||||
2. When an inbound peer is added, in this case the peer itself is set as the
|
||||
source of its own address
|
||||
3. When the switch is instructed to dial addresses via the `DialPeersAsync`
|
||||
method, in this case the node itself is set as the source
|
||||
|
||||
If the added address contains a node ID that is not registered in the address
|
||||
book, the address is added to a [bucket](#buckets) of new addresses.
|
||||
Otherwise, the additional address for an existing node ID is **not added** to
|
||||
the address book when:
|
||||
|
||||
- The last address added with the same node ID is stored in an old bucket, so
|
||||
it is considered a "good" address
|
||||
- There are addresses associated to the same node ID stored in
|
||||
`maxNewBucketsPerAddress = 4` distinct buckets
|
||||
- Randomly, with a probability that increases exponentially with the number of
|
||||
buckets in which there is an address with the same node ID.
|
||||
So, a new address for a node ID which is already present in one bucket is
|
||||
added with 50% of probability; if the node ID is present in two buckets, the
|
||||
probability decreases to 25%; and if it is present in three buckets, the
|
||||
probability is 12.5%.
|
||||
|
||||
The new address is also added to the `addrLookup` table, which stores
|
||||
`knownAddress` entries indexed by their node IDs.
|
||||
If the new address is from an unknown peer, a new entry is added to the
|
||||
`addrLookup` table; otherwise, the existing entry is updated with the new
|
||||
address.
|
||||
Entries of this table contain, among other fields, the list of buckets where
|
||||
addresses of a peer are stored.
|
||||
The `addrLookup` table is used by most of the address book methods (e.g.,
|
||||
`HasAddress`, `IsGood`, `MarkGood`, `MarkAttempt`), as it provides fast access
|
||||
to addresses.
|
||||
|
||||
### Errors
|
||||
|
||||
- if the added address or the associated source address are nil
|
||||
- if the added address is invalid
|
||||
- if the added address is the local node's address
|
||||
- if the added address ID is of a [banned](#bad-peers) peer
|
||||
- if either the added address or the associated source address IDs are configured as private IDs
|
||||
- if `routabilityStrict` is set and the address is not routable
|
||||
- in case of failures computing the bucket for the new address (`calcNewBucket` method)
|
||||
- if the added address instance, which is a new address, is configured as an
|
||||
old address (sanity check of `addToNewBucket` method)
|
||||
|
||||
## Need for Addresses
|
||||
|
||||
The `NeedMoreAddrs` method verifies whether the address book needs more addresses.
|
||||
|
||||
It is invoked by the PEX reactor to define whether to request peer addresses
|
||||
to a new outbound peer or to a randomly selected connected peer.
|
||||
|
||||
The address book needs more addresses when it has less than `1000` addresses
|
||||
registered, counting all buckets for new and old addresses.
|
||||
|
||||
## Pick address
|
||||
|
||||
The `PickAddress` method returns an address stored in the address book, chosen
|
||||
at random with a configurable bias towards new addresses.
|
||||
|
||||
It is invoked by the Peer Manager to obtain a peer address to dial, as part of
|
||||
its `ensurePeers` routine.
|
||||
The bias starts from 10%, when the peer has no outbound peers, increasing by
|
||||
10% for each outbound peer the node has, up to 90%, when the node has at least
|
||||
8 outbound peers.
|
||||
|
||||
The configured bias is a parameter that influences the probability of choosing
|
||||
an address from a bucket of new addresses or from a bucket of old addresses.
|
||||
A second parameter influencing this choice is the number of new and old
|
||||
addresses stored in the address book.
|
||||
In the absence of bias (i.e., if the configured bias is 50%), the probability
|
||||
of picking a new address is given by the square root of the number of new
|
||||
addresses divided by the sum of the square roots of the numbers of new and old
|
||||
addresses.
|
||||
By adding a bias toward new addresses (i.e., configured bias larger than 50%),
|
||||
the portion on the sample occupied by the square root of the number of new
|
||||
addresses increases, while the corresponding portion for old addresses decreases.
|
||||
As a result, it becomes more likely to pick a new address at random from this sample.
|
||||
|
||||
> The use of the square roots softens the impact of disproportional numbers of
|
||||
> new and old addresses in the address book. This is actually the expected
|
||||
> scenario, as there are 4 times more buckets for new addresses than buckets
|
||||
> for old addresses.
|
||||
|
||||
Once the type of address, new or old, is defined, a non-empty bucket of this
|
||||
type is selected at random.
|
||||
From the selected bucket, an address is chosen at random and returned.
|
||||
If all buckets of the selected type are empty, no address is returned.
|
||||
|
||||
## Random selection
|
||||
|
||||
The `GetSelection` method returns a selection of addresses stored in the
|
||||
address book, with no bias toward new or old addresses.
|
||||
|
||||
It is invoked by the PEX protocol to obtain a list of peer addresses with two
|
||||
purposes:
|
||||
|
||||
- To send to a peer in a PEX response, in the case of outbound peers or of
|
||||
nodes not operating in seed mode
|
||||
- To crawl, in the case of nodes operating in seed mode, as part of every
|
||||
interaction of the `crawlPeersRoutine`
|
||||
|
||||
The selection is a random subset of the peer addresses stored in the
|
||||
`addrLookup` table, which stores the last address added for each peer ID.
|
||||
The target size of the selection is `23%` (`getSelectionPercent`) of the
|
||||
number of addresses stored in the address book, but it should not be lower than
|
||||
`32` (`minGetSelection`) --- if it is, all addresses in the book are returned
|
||||
--- nor greater than `250` (`maxGetSelection`).
|
||||
|
||||
> The random selection is produced by:
|
||||
> - Retrieving all entries of the `addrLookup` map, which by definition are
|
||||
> returned in random order.
|
||||
> - Randomly shuffling the retrieved list, using the Fisher-Yates algorithm
|
||||
|
||||
## Random selection with bias
|
||||
|
||||
The `GetSelectionWithBias` method returns a selection of addresses stored in
|
||||
the address book, with bias toward new addresses.
|
||||
|
||||
It is invoked by the PEX protocol to obtain a list of peer addresses to be sent
|
||||
to a peer in a PEX response.
|
||||
This method is only invoked by seed nodes, when replying to a PEX request
|
||||
received from an inbound peer (i.e., a peer that dialed the seed node).
|
||||
The bias used in this scenario is hard-coded to 30%, meaning that 70% of
|
||||
the returned addresses are expected to be old addresses.
|
||||
|
||||
The number of addresses that compose the selection is computed in the same way
|
||||
as for the non-biased random selection.
|
||||
The bias toward new addresses is implemented by requiring that the configured
|
||||
bias, interpreted as a percentage, of the select addresses come from buckets of
|
||||
new addresses, while the remaining come from buckets of old addresses.
|
||||
Since the number of old addresses is typically lower than the number of new
|
||||
addresses, it is possible that the address book does not have enough old
|
||||
addresses to include in the selection.
|
||||
In this case, additional new addresses are included in the selection.
|
||||
Thus, the configured bias, in practice, is towards old addresses, not towards
|
||||
new addresses.
|
||||
|
||||
To randomly select addresses of a type, the address book considers all
|
||||
addresses present in every bucket of that type.
|
||||
This list of all addresses of a type is randomly shuffled, and the requested
|
||||
number of addresses are retrieved from the tail of this list.
|
||||
The returned selection contains, at its beginning, a random selection of new
|
||||
addresses in random order, followed by a random selection of old addresses, in
|
||||
random order.
|
||||
|
||||
## Dial Attempts
|
||||
|
||||
The `MarkAttempt` method records a failed attempt to connect to an address.
|
||||
|
||||
It is invoked by the Peer Manager when it fails dialing a peer, but the failure
|
||||
is not in the authentication step (`ErrSwitchAuthenticationFailure` error).
|
||||
In case of authentication errors, the peer is instead marked as a [bad peer](#bad-peers).
|
||||
|
||||
The failed connection attempt is recorded in the address registered for the
|
||||
peer's ID in the `addrLookup` table, which is the last address added with that ID.
|
||||
The known address' counter of failed `Attempts` is increased and the failure
|
||||
time is registered in `LastAttempt`.
|
||||
|
||||
The possible effect of recording multiple failed connect attempts to a peer is
|
||||
to turn its address into a *bad* address (do not confuse with banned addresses).
|
||||
A known address becomes bad if it is stored in buckets of new addresses, and
|
||||
when connection attempts:
|
||||
|
||||
- Have not been made over a week, i.e., `LastAttempt` is older than a week
|
||||
- Have failed 3 times and never succeeded, i.e., `LastSucess` field is unset
|
||||
- Have failed 10 times in the last week, i.e., `LastSucess` is older than a week
|
||||
|
||||
Addresses marked as *bad* are the first candidates to be removed from a bucket of
|
||||
new addresses when the bucket becomes full.
|
||||
|
||||
> Note that failed connection attempts are reported for a peer address, but in
|
||||
> fact the address book records them for a peer.
|
||||
>
|
||||
> More precisely, failed connection attempts are recorded in the entry of the
|
||||
> `addrLookup` table with reported peer ID, which contains the last address
|
||||
> added for that node ID, which is not necessarily the reported peer address.
|
||||
|
||||
## Good peers
|
||||
|
||||
The `MarkGood` method marks a peer ID as good.
|
||||
|
||||
It is invoked by the consensus reactor, via switch, when the number of useful
|
||||
messages received from a peer is a multiple of `10000`.
|
||||
Vote and block part messages are considered for this number, they must be valid
|
||||
and not be duplicated messages to be considered useful.
|
||||
|
||||
> The `SwitchReporter` type of `behaviour` package also invokes the `MarkGood`
|
||||
> method when a "reason" associated with consensus votes and block parts is
|
||||
> reported.
|
||||
> No reactor, however, currently provides these "reasons" to the `SwitchReporter`.
|
||||
|
||||
The effect of this action is that the address registered for the peer's ID in the
|
||||
`addrLookup` table, which is the last address added with that ID, is marked as
|
||||
good and moved to a bucket of old addresses.
|
||||
An address marked as good has its failed to connect counter and timestamp reset.
|
||||
If the destination bucket of old addresses is full, the oldest address in the
|
||||
bucket is moved (downgraded) to a bucket of new addresses.
|
||||
|
||||
Moving the peer address to a bucket of old addresses has the effect of
|
||||
upgrading, or increasing the ranking of a peer in the address book.
|
||||
|
||||
## Bad peers
|
||||
|
||||
The `MarkBad` method marks a peer as bad and bans it for a period of time.
|
||||
|
||||
This method is only invoked within the PEX reactor, with a banning time of 24
|
||||
hours, for the following reasons:
|
||||
|
||||
- A peer misbehaves in the [PEX protocol](pex-protocol.md#misbehavior)
|
||||
- When the `maxAttemptsToDial` limit (`16`) is reached for a peer
|
||||
- If an `ErrSwitchAuthenticationFailure` error is returned when dialing a peer
|
||||
|
||||
The effect of this action is that the address registered for the peer's ID in the
|
||||
`addrLookup` table, which is the last address added with that ID, is banned for
|
||||
a period of time.
|
||||
The banned peer is removed from the `addrLookup` table and from all buckets
|
||||
where its addresses are stored.
|
||||
|
||||
The information about banned peers, however, is not discarded.
|
||||
It is maintained in the `badPeers` map, indexed by peer ID.
|
||||
This allows, in particular, addresses of banned peers to be
|
||||
[reinstated](#reinstating-addresses), i.e., to be added
|
||||
back to the address book, when their ban period expires.
|
||||
|
||||
## Reinstating addresses
|
||||
|
||||
The `ReinstateBadPeers` method attempts to re-add banned addresses to the address book.
|
||||
|
||||
It is invoked by the PEX reactor when dialing new peers.
|
||||
This action is taken before requesting additional addresses to peers,
|
||||
in the case that the node needs more peer addresses.
|
||||
|
||||
The set of banned peer addresses is retrieved from the `badPeers` map.
|
||||
Addresses that are not any longer banned, i.e., whose banned period has expired,
|
||||
are added back to the address book as new addresses, while the corresponding
|
||||
node IDs are removed from the `badPeers` map.
|
||||
|
||||
## Removing addresses
|
||||
|
||||
The `RemoveAddress` method removes an address from the address book.
|
||||
|
||||
It is invoked by the switch when it dials a peer or accepts a connection from a
|
||||
peer that ends up being the node itself (`IsSelf` error).
|
||||
In both cases, the address dialed or accepted is also added to the address book
|
||||
as a local address, via the `AddOurAddress` method.
|
||||
|
||||
The same logic is also internally used by the address book for removing
|
||||
addresses of a peer that is [marked as a bad peer](#bad-peers).
|
||||
|
||||
The entry registered with the peer ID of the address in the `addrLookup` table,
|
||||
which is the last address added with that ID, is removed from all buckets where
|
||||
it is stored and from the `addrLookup` table.
|
||||
|
||||
> FIXME: is it possible that addresses with the same ID as the removed address,
|
||||
> but with distinct network addresses, are kept in buckets of the address book?
|
||||
> While they will not be accessible anymore, as there is no reference to them
|
||||
> in the `addrLookup`, they will still be there.
|
||||
|
||||
## Persistence
|
||||
|
||||
The `loadFromFile` method, called when the address book is started, reads
|
||||
address book entries from a file, passed to the address book constructor.
|
||||
The file, at this point, does not need to exist.
|
||||
|
||||
The `saveRoutine` is started when the address book is started.
|
||||
It saves the address book to the configured file every `dumpAddressInterval`,
|
||||
hard-coded to 2 minutes.
|
||||
It is also possible to save the content of the address book using the `Save`
|
||||
method.
|
||||
Saving the address book content to a file acquires the address book lock, also
|
||||
employed by all other public methods.
|
||||
@@ -0,0 +1,51 @@
|
||||
# Tendermint p2p configuration
|
||||
|
||||
This document contains configurable parameters a node operator can use to tune the p2p behaviour.
|
||||
|
||||
| Parameter| Default| Description |
|
||||
| --- | --- | ---|
|
||||
| ListenAddress | "tcp://0.0.0.0:26656" | Address to listen for incoming connections (0.0.0.0:0 means any interface, any port) |
|
||||
| ExternalAddress | "" | Address to advertise to peers for them to dial |
|
||||
| [Seeds](pex-protocol.md#seed-nodes) | empty | Comma separated list of seed nodes to connect to (ID@host:port )|
|
||||
| [Persistent peers](peer_manager.md#persistent-peers) | empty | Comma separated list of nodes to keep persistent connections to (ID@host:port ) |
|
||||
| UPNP | false | UPNP port forwarding enabled |
|
||||
| [AddrBook](addressbook.md) | defaultAddrBookPath | Path do address book |
|
||||
| AddrBookStrict | true | Set true for strict address routability rules and false for private or local networks |
|
||||
| [MaxNumInboundPeers](switch.md#accepting-peers) | 40 | Maximum number of inbound peers |
|
||||
| [MaxNumOutboundPeers](peer_manager.md#ensure-peers) | 10 | Maximum number of outbound peers to connect to, excluding persistent peers |
|
||||
| [UnconditionalPeers](switch.md#accepting-peers) | empty | These are IDs of the peers which are allowed to be (re)connected as both inbound or outbound regardless of whether the node reached `max_num_inbound_peers` or `max_num_outbound_peers` or not. |
|
||||
| PersistentPeersMaxDialPeriod| 0 * time.Second | Maximum pause when redialing a persistent peer (if zero, exponential backoff is used) |
|
||||
| FlushThrottleTimeout |100 * time.Millisecond| Time to wait before flushing messages out on the connection |
|
||||
| MaxPacketMsgPayloadSize | 1024 | Maximum size of a message packet payload, in bytes |
|
||||
| SendRate | 5120000 (5 mB/s) | Rate at which packets can be sent, in bytes/second |
|
||||
| RecvRate | 5120000 (5 mB/s) | Rate at which packets can be received, in bytes/second|
|
||||
| [PexReactor](pex.md) | true | Set true to enable the peer-exchange reactor |
|
||||
| SeedMode | false | Seed mode, in which node constantly crawls the network and looks for. Does not work if the peer-exchange reactor is disabled. |
|
||||
| PrivatePeerIDs | empty | Comma separated list of peer IDsthat we do not add to the address book or gossip to other peers. They stay private to us. |
|
||||
| AllowDuplicateIP | false | Toggle to disable guard against peers connecting from the same ip.|
|
||||
| [HandshakeTimeout](transport.md#connection-upgrade) | 20 * time.Second | Timeout for handshake completion between peers |
|
||||
| [DialTimeout](switch.md#dialing-peers) | 3 * time.Second | Timeout for dialing a peer |
|
||||
|
||||
|
||||
These parameters can be set using the `$TMHOME/config/config.toml` file. A subset of them can also be changed via command line using the following command line flags:
|
||||
|
||||
| Parameter | Flag| Example|
|
||||
| --- | --- | ---|
|
||||
| Listen address| `p2p.laddr` | "tcp://0.0.0.0:26656" |
|
||||
| Seed nodes | `p2p.seeds` | `--p2p.seeds “id100000000000000000000000000000000@1.2.3.4:26656,id200000000000000000000000000000000@2.3.4.5:4444”` |
|
||||
| Persistent peers | `p2p.persistent_peers` | `--p2p.persistent_peers “id100000000000000000000000000000000@1.2.3.4:26656,id200000000000000000000000000000000@2.3.4.5:26656”` |
|
||||
| Unconditional peers | `p2p.unconditional_peer_ids` | `--p2p.unconditional_peer_ids “id100000000000000000000000000000000,id200000000000000000000000000000000”` |
|
||||
| UPNP | `p2p.upnp` | `--p2p.upnp` |
|
||||
| PexReactor | `p2p.pex` | `--p2p.pex` |
|
||||
| Seed mode | `p2p.seed_mode` | `--p2p.seed_mode` |
|
||||
| Private peer ids | `p2p.private_peer_ids` | `--p2p.private_peer_ids “id100000000000000000000000000000000,id200000000000000000000000000000000”` |
|
||||
|
||||
**Note on persistent peers**
|
||||
|
||||
If `persistent_peers_max_dial_period` is set greater than zero, the
|
||||
pause between each dial to each persistent peer will not exceed `persistent_peers_max_dial_period`
|
||||
during exponential backoff and we keep trying again without giving up.
|
||||
|
||||
If `seeds` and `persistent_peers` intersect,
|
||||
the user will be warned that seeds may auto-close connections
|
||||
and that the node may not be able to keep the connection persistent.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 129 KiB |
@@ -0,0 +1,147 @@
|
||||
# Peer Manager
|
||||
|
||||
The peer manager is responsible for establishing connections with peers.
|
||||
It defines when a node should dial peers and which peers it should dial.
|
||||
The peer manager is not an implementation abstraction of the p2p layer,
|
||||
but a role that is played by the [PEX reactor](./pex.md).
|
||||
|
||||
## Outbound peers
|
||||
|
||||
The `ensurePeersRoutine` is a persistent routine intended to ensure that a node
|
||||
is connected to `MaxNumOutboundPeers` outbound peers.
|
||||
This routine is continuously executed by regular nodes, i.e. nodes not
|
||||
operating in seed mode, as part of the PEX reactor implementation.
|
||||
|
||||
The logic defining when the node should dial peers, for selecting peers to dial
|
||||
and for actually dialing them is implemented in the `ensurePeers` method.
|
||||
This method is periodically invoked -- every `ensurePeersPeriod`, with default
|
||||
value to 30 seconds -- by the `ensurePeersRoutine`.
|
||||
|
||||
A node is expected to dial peers whenever the number of outbound peers is lower
|
||||
than the configured `MaxNumOutboundPeers` parameter.
|
||||
The current number of outbound peers is retrieved from the switch, using the
|
||||
`NumPeers` method, which also reports the number of nodes to which the switch
|
||||
is currently dialing.
|
||||
If the number of outbound peers plus the number of dialing routines equals to
|
||||
`MaxNumOutboundPeers`, nothing is done.
|
||||
Otherwise, the `ensurePeers` method will attempt to dial node addresses in
|
||||
order to reach the target number of outbound peers.
|
||||
|
||||
Once defined that the node needs additional outbound peers, the node queries
|
||||
the address book for candidate addresses.
|
||||
This is done using the [`PickAddress`](./addressbook.md#pick-address) method,
|
||||
which returns an address selected at random on the address book, with some bias
|
||||
towards new or old addresses.
|
||||
When the node has up to 3 outbound peers, the adopted bias is towards old
|
||||
addresses, i.e., addresses of peers that are believed to be "good".
|
||||
When the node has from 5 outbound peers, the adopted bias is towards new
|
||||
addresses, i.e., addresses of peers about which the node has not yet collected
|
||||
much information.
|
||||
So, the more outbound peers a node has, the less conservative it will be when
|
||||
selecting new peers.
|
||||
|
||||
The selected peer addresses are then dialed in parallel, by starting a dialing
|
||||
routine per peer address.
|
||||
Dialing a peer address can fail for multiple reasons.
|
||||
The node might have attempted to dial the peer too many times.
|
||||
In this case, the peer address is marked as bad and removed from the address book.
|
||||
The node might have attempted and failed to dial the peer recently
|
||||
and the exponential `backoffDuration` has not yet passed.
|
||||
Or the current connection attempt might fail, which is registered in the address book.
|
||||
None of these errors are explicitly handled by the `ensurePeers` method, which
|
||||
also does not wait until the connections are established.
|
||||
|
||||
The third step of the `ensurePeers` method is to ensure that the address book
|
||||
has enough addresses.
|
||||
This is done, first, by [reinstating banned peers](./addressbook.md#Reinstating-addresses)
|
||||
whose ban period has expired.
|
||||
Then, the node randomly selects a connected peer, which can be either an
|
||||
inbound or outbound peer, to [requests addresses](./pex-protocol.md#Requesting-Addresses)
|
||||
using the PEX protocol.
|
||||
Last, and this action is only performed if the node could not retrieve any new
|
||||
address to dial from the address book, the node dials the configured seed nodes
|
||||
in order to establish a connection to at least one of them.
|
||||
|
||||
### Fast dialing
|
||||
|
||||
As above described, seed nodes are actually the last source of peer addresses
|
||||
for regular nodes.
|
||||
They are contacted by a node when, after an invocation of the `ensurePeers`
|
||||
method, no suitable peer address to dial is retrieved from the address book
|
||||
(e.g., because it is empty).
|
||||
|
||||
Once a connection with a seed node is established, the node immediately
|
||||
[sends a PEX request](./pex-protocol.md#Requesting-Addresses) to it, as it is
|
||||
added as an outbound peer.
|
||||
When the corresponding PEX response is received, the addresses provided by the
|
||||
seed node are added to the address book.
|
||||
As a result, in the next invocation of the `ensurePeers` method, the node
|
||||
should be able to dial some of the peer addresses provided by the seed node.
|
||||
|
||||
However, as observed in this [issue](https://github.com/tendermint/tendermint/issues/2093),
|
||||
it can take some time, up to `ensurePeersPeriod` or 30 seconds, from when the
|
||||
node receives new peer addresses and when it dials the received addresses.
|
||||
To avoid this delay, which can be particularly relevant when the node has no
|
||||
peers, a node immediately attempts to dial peer addresses when they are
|
||||
received from a peer that is locally configured as a seed node.
|
||||
|
||||
> FIXME: The current logic was introduced in [#3762](https://github.com/tendermint/tendermint/pull/3762).
|
||||
> Although it fix the issue, the delay between receiving an address and dialing
|
||||
> the peer, it does not impose and limit on how many addresses are dialed in this
|
||||
> scenario.
|
||||
> So, all addresses received from a seed node are dialed, regardless of the
|
||||
> current number of outbound peers, the number of dialing routines, or the
|
||||
> `MaxNumOutboundPeers` parameter.
|
||||
>
|
||||
> Issue [#9548](https://github.com/tendermint/tendermint/issues/9548) was
|
||||
> created to handle this situation.
|
||||
|
||||
### First round
|
||||
|
||||
When the PEX reactor is started, the `ensurePeersRoutine` is created and it
|
||||
runs thorough the operation of a node, periodically invoking the `ensurePeers`
|
||||
method.
|
||||
However, if when the persistent routine is started the node already has some
|
||||
peers, either inbound or outbound peers, or is dialing some addresses, the
|
||||
first invocation of `ensurePeers` is delayed by a random amount of time from 0
|
||||
to `ensurePeersPeriod`.
|
||||
|
||||
### Persistent peers
|
||||
|
||||
The node configuration can contain a list of *persistent peers*.
|
||||
Those peers have preferential treatment compared to regular peers and the node
|
||||
is always trying to connect to them.
|
||||
Moreover, these peers are not removed from the address book in the case of
|
||||
multiple failed dial attempts.
|
||||
|
||||
On startup, the node immediately tries to dial the configured persistent peers
|
||||
by calling the switch's [`DialPeersAsync`](./switch.md#manual-operation) method.
|
||||
This is not done in the p2p package, but it is part of the procedure to set up a node.
|
||||
|
||||
> TODO: the handling of persistent peers should be described in more detail.
|
||||
|
||||
### Life cycle
|
||||
|
||||
The picture below is a first attempt of illustrating the life cycle of an outbound peer:
|
||||
|
||||
<img src="img/p2p_state.png" width="50%" title="Outgoing peers lifecycle">
|
||||
|
||||
A peer can be in the following states:
|
||||
|
||||
- Candidate peers: peer addresses stored in the address boook, that can be
|
||||
retrieved via the [`PickAddress`](./addressbook.md#pick-address) method
|
||||
- [Dialing](switch.md#dialing-peers): peer addresses that are currently being
|
||||
dialed. This state exists to ensure that a single dialing routine exist per peer.
|
||||
- [Reconnecting](switch.md#reconnect-to-peer): persistent peers to which a node
|
||||
is currently reconnecting, as a previous connection attempt has failed.
|
||||
- Connected peers: peers that a node has successfully dialed, added as outbound peers.
|
||||
- [Bad peers](addressbook.md#bad-peers): peers marked as bad in the address
|
||||
book due to exhibited [misbehavior](pex-protocol.md#misbehavior).
|
||||
Peers can be reinstated after being marked as bad.
|
||||
|
||||
## Pending of documentation
|
||||
|
||||
The `dialSeeds` method of the PEX reactor.
|
||||
|
||||
The `dialPeer` method of the PEX reactor.
|
||||
This includes `dialAttemptsInfo`, `maxBackoffDurationForPeer` methods.
|
||||
@@ -0,0 +1,240 @@
|
||||
# Peer Exchange Protocol
|
||||
|
||||
The Peer Exchange (PEX) protocol enables nodes to exchange peer addresses, thus
|
||||
implementing a peer discovery mechanism.
|
||||
|
||||
The PEX protocol uses two messages:
|
||||
|
||||
- `PexRequest`: sent by a node to [request](#requesting-addresses) peer
|
||||
addresses to a peer
|
||||
- `PexAddrs`: a list of peer addresses [provided](#providing-addresses) to a
|
||||
peer as response to a `PexRequest` message
|
||||
|
||||
While all nodes, with few exceptions, participate on the PEX protocol,
|
||||
a subset of nodes, configured as [seed nodes](#seed-nodes) have a particular
|
||||
role in the protocol.
|
||||
They crawl the network, connecting to random peers, in order to learn as many
|
||||
peer addresses as possible to provide to other nodes.
|
||||
|
||||
## Requesting Addresses
|
||||
|
||||
A node requests peer addresses by sending a `PexRequest` message to a peer.
|
||||
|
||||
For regular nodes, not operating in seed mode, a PEX request is sent when
|
||||
the node *needs* peers addresses, a condition checked:
|
||||
|
||||
1. When an *outbound* peer is added, causing the node to request addresses from
|
||||
the new peer
|
||||
2. Periodically, by the `ensurePeersRoutine`, causing the node to request peer
|
||||
addresses to a randomly selected peer
|
||||
|
||||
A node needs more peer addresses when its addresses book has
|
||||
[less than 1000 records](./addressbook.md#need-for-addresses).
|
||||
It is thus reasonable to assume that the common case is that a peer needs more
|
||||
peer addresses, so that PEX requests are sent whenever the above two situations happen.
|
||||
|
||||
A PEX request is sent when a new *outbound* peer is added.
|
||||
The same does not happen with new inbound peers because the implementation
|
||||
considers outbound peers, that the node has chosen for dialing, more
|
||||
trustworthy than inbound peers, that the node has accepted.
|
||||
Moreover, when a node is short of peer addresses, it dials the configured seed nodes;
|
||||
since they are added as outbound peers, the node can immediately request peer addresses.
|
||||
|
||||
The `ensurePeersRoutine` periodically checks, by default every 30 seconds (`ensurePeersPeriod`),
|
||||
whether the node has enough outbound peers.
|
||||
If it does not have, the node tries dialing some peer addresses stored in the address book.
|
||||
As part of this procedure, the node selects a peer at random,
|
||||
from the set of connected peers retrieved from the switch,
|
||||
and sends a PEX request to the selected peer.
|
||||
|
||||
Sending a PEX request to a peer is implemented by the `RequestAddrs` method of
|
||||
the PEX reactor.
|
||||
|
||||
### Responses
|
||||
|
||||
After a PEX request is sent to a peer, the node expects to receive,
|
||||
as a response, a `PexAddrs` message from the peer.
|
||||
This message encodes a list of peer addresses that are
|
||||
[added to address book](./addressbook.md#adding-addresses),
|
||||
having the peer from which the PEX response was received as their source.
|
||||
|
||||
Received PEX responses are handled by the `ReceiveAddrs` method of the PEX reactor.
|
||||
In the case of a PEX response received from a peer which is configured as
|
||||
a seed node, the PEX reactor attempts immediately to dial the provided peer
|
||||
addresses, as detailed [here](./peer_manager.md#fast-dialing).
|
||||
|
||||
### Misbehavior
|
||||
|
||||
Sending multiple PEX requests to a peer, before receiving a reply from it,
|
||||
is considered a misbehavior.
|
||||
To prevent it, the node maintains a `requestsSent` set of outstanding
|
||||
requests, indexed by destination peers.
|
||||
While a peer ID is present in the `requestsSent` set, the node does not send
|
||||
further PEX requests to that peer.
|
||||
A peer ID is removed from the `requestsSent` set when a PEX response is
|
||||
received from it.
|
||||
|
||||
Sending a PEX response to a peer that has not requested peer addresses
|
||||
is also considered a misbehavior.
|
||||
So, if a PEX response is received from a peer that is not registered in
|
||||
the `requestsSent` set, a `ErrUnsolicitedList` error is produced.
|
||||
This leads the peer to be disconnected and [marked as a bad peer](addressbook.md#bad-peers).
|
||||
|
||||
## Providing Addresses
|
||||
|
||||
When a node receives a `PexRequest` message from a peer,
|
||||
it replies with a `PexAddrs` message.
|
||||
|
||||
This message encodes a [random selection of peer addresses](./addressbook.md#random-selection)
|
||||
retrieved from the address book.
|
||||
|
||||
Sending a PEX response to a peer is implemented by the `SendAddrs` method of
|
||||
the PEX reactor.
|
||||
|
||||
### Misbehavior
|
||||
|
||||
Requesting peer addresses too often is considered a misbehavior.
|
||||
Since node are expected to send PEX requests every `ensurePeersPeriod`,
|
||||
the minimum accepted interval between requests from the same peer is set
|
||||
to `ensurePeersPeriod / 3`, 10 seconds by default.
|
||||
|
||||
The `receiveRequest` method is responsible for verifying this condition.
|
||||
The node keeps a `lastReceivedRequests` map with the time of the last PEX
|
||||
request received from every peer.
|
||||
If the interval between successive requests is less than the minimum accepted
|
||||
one, the peer is disconnected and [marked as a bad peer](addressbook.md#bad-peers).
|
||||
An exception is made for the first two PEX requests received from a peer.
|
||||
|
||||
> The probably reason is that, when a new peer is added, the two conditions for
|
||||
> a node to request peer addresses can be triggered with an interval lower than
|
||||
> the minimum accepted interval.
|
||||
> Since this is a legit behavior, it should not be punished.
|
||||
|
||||
## Seed nodes
|
||||
|
||||
A seed node is a node configured to operate in `SeedMode`.
|
||||
|
||||
### Crawling peers
|
||||
|
||||
Seed nodes crawl the network, connecting to random peers and sending PEX
|
||||
requests to them, in order to learn as many peer addresses as possible.
|
||||
More specifically, a node operating in seed mode sends PEX requests in two cases:
|
||||
|
||||
1. When an outbound peer is added, and the seed node needs more peer addresses,
|
||||
it requests peer addresses to the new peer
|
||||
2. Periodically, the `crawlPeersRoutine` sends PEX requests to a random set of
|
||||
peers, whose addresses are registered in the Address Book
|
||||
|
||||
The first case also applies for nodes not operating in seed mode.
|
||||
The second case replaces the second for regular nodes, as seed nodes do not
|
||||
run the `ensurePeersRoutine`, as regular nodes,
|
||||
but run the `crawlPeersRoutine`, which is not run by regular nodes.
|
||||
|
||||
The `crawlPeersRoutine` periodically, every 30 seconds (`crawlPeerPeriod`),
|
||||
starts a new peer discovery round.
|
||||
First, the seed node retrieves a random selection of peer addresses from its
|
||||
Address Book.
|
||||
This selection is produced in the same way as in the random selection of peer
|
||||
addresses that are [provided](#providing-addresses) to a requesting peer.
|
||||
Peers that the seed node has crawled recently,
|
||||
less than 2 minutes ago (`minTimeBetweenCrawls`), are removed from this selection.
|
||||
The remaining peer addresses are registered in the `crawlPeerInfos` table.
|
||||
|
||||
The seed node is not necessarily connected to the peer whose address is
|
||||
selected for each round of crawling.
|
||||
So, the seed node dials the selected peer addresses.
|
||||
This is performed in foreground, one peer at a time.
|
||||
As a result, a round of crawling can take a substantial amount of time.
|
||||
For each selected peer it succeeds dialing to, this include already connected
|
||||
peers, the seed node sends a PEX request.
|
||||
|
||||
Dialing a selected peer address can fail for multiple reasons.
|
||||
The seed node might have attempted to dial the peer too many times.
|
||||
In this case, the peer address is marked as [bad in the address book](addressbook.md#bad-peers).
|
||||
The seed node might have attempted to dial the peer recently, without success,
|
||||
and the exponential `backoffDuration` has not yet passed.
|
||||
Or the current connection attempt might fail, which is registered in the address book.
|
||||
|
||||
Failures to dial to a peer address produce an information that is important for
|
||||
a seed node.
|
||||
They indicate that a peer is unreachable, or is not operating correctly, and
|
||||
therefore its address should not be provided to other nodes.
|
||||
This occurs when, due to multiple failed connection attempts or authentication
|
||||
failures, the peer address ends up being removed from the address book.
|
||||
As a result, the periodically crawling of selected peers not only enables the
|
||||
discovery of new peers, but also allows the seed node to stop providing
|
||||
addresses of bad peers.
|
||||
|
||||
### Offering addresses
|
||||
|
||||
Nodes operating in seed mode handle PEX requests differently than regular
|
||||
nodes, whose operation is described [here](#providing-addresses).
|
||||
|
||||
This distinction exists because nodes dial a seed node with the main, if not
|
||||
exclusive goal of retrieving peer addresses.
|
||||
In other words, nodes do not dial a seed node because they intend to have it as
|
||||
a peer in the multiple Tendermint protocols, but because they believe that a
|
||||
seed node is a good source of addresses of nodes to which they can establish
|
||||
connections and interact in the multiple Tendermint protocols.
|
||||
|
||||
So, when a seed node receives a `PexRequest` message from an inbound peer,
|
||||
it sends a `PexAddrs` message, containing a selection of peer
|
||||
addresses, back to the peer and *disconnects* from it.
|
||||
Seed nodes therefore treat inbound connections from peers as a short-term
|
||||
connections, exclusively intended to retrieve peer addresses.
|
||||
Once the requested peer addresses are sent, the connection with the peer is closed.
|
||||
|
||||
Moreover, the selection of peer addresses provided to inbound peers by a seed
|
||||
node, although still essentially random, has a [bias toward old
|
||||
addresses](./addressbook.md#random-selection-with-bias).
|
||||
The selection bias is defined by `biasToSelectNewPeers`, hard-coded to `30%`,
|
||||
meaning that `70%` of the peer addresses provided by a seed node are expected
|
||||
to be old addresses.
|
||||
Although this nomenclature is not clear, *old* addresses are the addresses that
|
||||
survived the most in the address book, that is, are addresses that the seed
|
||||
node believes being from *good* peers (more details [here](./addressbook.md#good-peers)).
|
||||
|
||||
Another distinction is on the handling of potential [misbehavior](#misbehavior-1)
|
||||
of peers requesting addresses.
|
||||
A seed node does not enforce, a priori, a minimal interval between PEX requests
|
||||
from inbound peers.
|
||||
Instead, it does not reply to more than one PEX request per peer inbound
|
||||
connection, and, as above mentioned, it disconnects from incoming peers after
|
||||
responding to them.
|
||||
If the same peer dials again to the seed node and requests peer addresses, the
|
||||
seed node will reply to this peer like it was the first time it has requested
|
||||
peer addresses.
|
||||
|
||||
> This is more an implementation restriction than a desired behavior.
|
||||
> The `lastReceivedRequests` map stores the last time a PEX request was
|
||||
> received from a peer, and the entry relative to a peer is removed from this
|
||||
> map when the peer is disconnected.
|
||||
>
|
||||
> It is debatable whether this approach indeed prevents abuse against seed nodes.
|
||||
|
||||
### Disconnecting from peers
|
||||
|
||||
Seed nodes treat connections with peers as short-term connections, which are
|
||||
mainly, if not exclusively, intended to exchange peer addresses.
|
||||
|
||||
In the case of inbound peers, that have dialed the seed node, the intent of the
|
||||
connection is achieved once a PEX response is sent to the peer.
|
||||
The seed node thus disconnects from an inbound peer after sending a `PexAddrs`
|
||||
message to it.
|
||||
|
||||
In the case of outbound peers, which the seed node has dialed for crawling peer
|
||||
addresses, the intent of the connection is essentially achieved when a PEX
|
||||
response is received from the peer.
|
||||
The seed node, however, does not disconnect from a peer after receiving a
|
||||
selection of peer addresses from it.
|
||||
As a result, after some rounds of crawling, a seed node will have established
|
||||
connections to a substantial amount of peers.
|
||||
|
||||
To couple with the existence of multiple connections with peers that have no
|
||||
longer purpose for the seed node, the `crawlPeersRoutine` also invokes, after
|
||||
each round of crawling, the `attemptDisconnects` method.
|
||||
This method retrieves the list of connected peers from the switch, and
|
||||
disconnects from peers that are not persistent peers, and with which a
|
||||
connection is established for more than `SeedDisconnectWaitPeriod`.
|
||||
This period is a configuration parameter, set to 28 hours when the PEX reactor
|
||||
is created by the default node constructor.
|
||||
@@ -0,0 +1,111 @@
|
||||
# PEX Reactor
|
||||
|
||||
The PEX reactor is one of the reactors running in a Tendermint node.
|
||||
|
||||
Its implementation is located in the `p2p/pex` package, and it is considered
|
||||
part of the implementation of the p2p layer.
|
||||
|
||||
This document overviews the implementation of the PEX reactor, describing how
|
||||
the methods from the `Reactor` interface are implemented.
|
||||
|
||||
The actual operation of the PEX reactor is presented in documents describing
|
||||
the roles played by the PEX reactor in the p2p layer:
|
||||
|
||||
- [Address Book](./addressbook.md): stores known peer addresses and information
|
||||
about peers to which the node is connected or has attempted to connect
|
||||
- [Peer Manager](./peer_manager.md): manages connections established with peers,
|
||||
defining when a node should dial peers and which peers it should dial
|
||||
- [Peer Exchange protocol](./pex-protocol.md): enables nodes to exchange peer
|
||||
addresses, thus implementing a peer discovery service
|
||||
|
||||
## OnStart
|
||||
|
||||
The `OnStart` method implements `BaseService` and starts the PEX reactor.
|
||||
|
||||
The [address book](./addressbook.md), which is a `Service` is started.
|
||||
This loads the address book content from disk,
|
||||
and starts a routine that periodically persists the address book content to disk.
|
||||
|
||||
The PEX reactor is configured with the addresses of a number of seed nodes,
|
||||
the `Seeds` parameter of the `ReactorConfig`.
|
||||
The addresses of seed nodes are parsed into `NetAddress` instances and resolved
|
||||
into IP addresses, which is implemented by the `checkSeeds` method.
|
||||
Valid seed node addresses are stored in the `seedAddrs` field,
|
||||
and are used by the `dialSeeds` method to contact the configured seed nodes.
|
||||
|
||||
The last action is to start one of the following persistent routines, based on
|
||||
the `SeedMode` configuration parameter:
|
||||
|
||||
- Regular nodes run the `ensurePeersRoutine` to check whether the node has
|
||||
enough outbound peers, dialing peers when necessary
|
||||
- Seed nodes run the `crawlPeersRoutine` to periodically start a new round
|
||||
of [crawling](./pex-protocol.md#Crawling-peers) to discover as many peer
|
||||
addresses as possible
|
||||
|
||||
### Errors
|
||||
|
||||
Errors encountered when loading the address book from disk are returned,
|
||||
and prevent the reactor from being started.
|
||||
An exception is made for the `service.ErrAlreadyStarted` error, which is ignored.
|
||||
|
||||
Errors encountered when parsing the configured addresses of seed nodes
|
||||
are returned and cause the reactor startup to fail.
|
||||
An exception is made for DNS resolution `ErrNetAddressLookup` errors,
|
||||
which are not deemed fatal and are only logged as invalid addresses.
|
||||
|
||||
If none of the configured seed node addresses is valid, and the loaded address
|
||||
book is empty, the reactor is not started and an error is returned.
|
||||
|
||||
## OnStop
|
||||
|
||||
The `OnStop` method implements `BaseService` and stops the PEX reactor.
|
||||
|
||||
The address book routine that periodically saves its content to disk is stopped.
|
||||
|
||||
## GetChannels
|
||||
|
||||
The `GetChannels` method, from the `Reactor` interface, returns the descriptor
|
||||
of the channel used by the PEX protocol.
|
||||
|
||||
The channel ID is `PexChannel` (0), with priority `1`, send queue capacity of
|
||||
`10`, and maximum message size of `64000` bytes.
|
||||
|
||||
## AddPeer
|
||||
|
||||
The `AddPeer` method, from the `Reactor` interface,
|
||||
adds a new peer to the PEX protocol.
|
||||
|
||||
If the new peer is an **inbound peer**, i.e., if the peer has dialed the node,
|
||||
the peer's address is [added to the address book](./addressbook.md#adding-addresses).
|
||||
Since the peer was authenticated when establishing a secret connection with it,
|
||||
the source of the peer address is trusted, and its source is set by the peer itself.
|
||||
In the case of an outbound peer, the node should already have its address in
|
||||
the address book, as the switch has dialed the peer.
|
||||
|
||||
If the peer is an **outbound peer**, i.e., if the node has dialed the peer,
|
||||
and the PEX protocol needs more addresses,
|
||||
the node [sends a PEX request](./pex-protocol.md#Requesting-Addresses) to the peer.
|
||||
The same is not done when inbound peers are added because they are deemed least
|
||||
trustworthy than outbound peers.
|
||||
|
||||
## RemovePeer
|
||||
|
||||
The `RemovePeer` method, from the `Reactor` interface,
|
||||
removes a peer from the PEX protocol.
|
||||
|
||||
The peer's ID is removed from the tables tracking PEX requests
|
||||
[sent](./pex-protocol.md#misbehavior) but not yet replied
|
||||
and PEX requests [received](./pex-protocol.md#misbehavior-1).
|
||||
|
||||
## Receive
|
||||
|
||||
The `Receive` method, from the `Reactor` interface,
|
||||
handles a message received by the PEX protocol.
|
||||
|
||||
A node receives two type of messages as part of the PEX protocol:
|
||||
|
||||
- `PexRequest`: a request for addresses received from a peer, handled as
|
||||
described [here](./pex-protocol.md#providing-addresses)
|
||||
- `PexAddrs`: a list of addresses received from a peer, as a reponse to a PEX
|
||||
request sent by the node, as described [here](./pex-protocol.md#responses)
|
||||
|
||||
@@ -0,0 +1,237 @@
|
||||
# Switch
|
||||
|
||||
The switch is a core component of the p2p layer.
|
||||
It manages the procedures for [dialing peers](#dialing-peers) and
|
||||
[accepting](#accepting-peers) connections from peers, which are actually
|
||||
implemented by the [transport](./transport.md).
|
||||
It also manages the reactors, i.e., protocols implemented by the node that
|
||||
interact with its peers.
|
||||
Once a connection with a peer is established, the peer is [added](#add-peer) to
|
||||
the switch and all registered reactors.
|
||||
Reactors may also instruct the switch to [stop a peer](#stop-peer), namely
|
||||
disconnect from it.
|
||||
The switch, in this case, makes sure that the peer is removed from all
|
||||
registered reactors.
|
||||
|
||||
## Dialing peers
|
||||
|
||||
Dialing a peer is implemented by the `DialPeerWithAddress` method.
|
||||
|
||||
This method is invoked by the [peer manager](./peer_manager.md#ensure-peers)
|
||||
to dial a peer address and establish a connection with an outbound peer.
|
||||
|
||||
The switch keeps a single dialing routine per peer ID.
|
||||
This is ensured by keeping a synchronized map `dialing` with the IDs of peers
|
||||
to which the peer is dialing.
|
||||
A peer ID is added to `dialing` when the `DialPeerWithAddress` method is called
|
||||
for that peer, and it is removed when the method returns for whatever reason.
|
||||
The method returns immediately when invoked for a peer which ID is already in
|
||||
the `dialing` structure.
|
||||
|
||||
The actual dialing is implemented by the [`Dial`](./transport.md#dial) method
|
||||
of the transport configured for the switch, in the `addOutboundPeerWithConfig`
|
||||
method.
|
||||
If the transport succeeds establishing a connection, the returned `Peer` is
|
||||
added to the switch using the [`addPeer`](#add-peer) method.
|
||||
This operation can fail, returning an error. In this case, the switch invokes
|
||||
the transport's [`Cleanup`](./transport.md#cleanup) method to clean any resources
|
||||
associated with the peer.
|
||||
|
||||
If the transport fails to establish a connection with the peer that is configured
|
||||
as a persistent peer, the switch spawns a routine to [reconnect to the peer](#reconnect-to-peer).
|
||||
If the peer is already in the `reconnecting` state, the spawned routine has no
|
||||
effect and returns immediately.
|
||||
This is in fact a likely scenario, as the `reconnectToPeer` routine relies on
|
||||
this same `DialPeerWithAddress` method for dialing peers.
|
||||
|
||||
### Manual operation
|
||||
|
||||
The `DialPeersAsync` method receives a list of peer addresses (strings)
|
||||
and dials all of them in parallel.
|
||||
It is invoked in two situations:
|
||||
|
||||
- In the [setup](https://github.com/tendermint/tendermint/blob/29c5a062d23aaef653f11195db55c45cd9e02715/node/node.go#L985) of a node, to establish connections with every configured
|
||||
persistent peer
|
||||
- In the RPC package, to implement two unsafe RPC commands, not used in production:
|
||||
[`DialSeeds`](https://github.com/tendermint/tendermint/blob/29c5a062d23aaef653f11195db55c45cd9e02715/rpc/core/net.go#L47) and
|
||||
[`DialPeers`](https://github.com/tendermint/tendermint/blob/29c5a062d23aaef653f11195db55c45cd9e02715/rpc/core/net.go#L87)
|
||||
|
||||
The received list of peer addresses to dial is parsed into `NetAddress` instances.
|
||||
In case of parsing errors, the method returns. An exception is made for
|
||||
DNS resolution `ErrNetAddressLookup` errors, which do not interrupt the procedure.
|
||||
|
||||
As the peer addresses provided to this method are typically not known by the node,
|
||||
contrarily to the addressed dialed using the `DialPeerWithAddress` method,
|
||||
they are added to the node's address book, which is persisted to disk.
|
||||
|
||||
The switch dials the provided peers in parallel.
|
||||
The list of peer addresses is randomly shuffled, and for each peer a routine is
|
||||
spawned.
|
||||
Each routine sleeps for a random interval, up to 3 seconds, then invokes the
|
||||
`DialPeerWithAddress` method that actually dials the peer.
|
||||
|
||||
### Reconnect to peer
|
||||
|
||||
The `reconnectToPeer` method is invoked when a connection attempt to a peer fails,
|
||||
and the peer is configured as a persistent peer.
|
||||
|
||||
The `reconnecting` synchronized map keeps the peer's in this state, identified
|
||||
by their IDs (string).
|
||||
This should ensure that a single instance of this method is running at any time.
|
||||
The peer is kept in this map while this method is running for it: it is set on
|
||||
the beginning, and removed when the method returns for whatever reason.
|
||||
If the peer is already in the `reconnecting` state, nothing is done.
|
||||
|
||||
The remaining of the method performs multiple connection attempts to the peer,
|
||||
via `DialPeerWithAddress` method.
|
||||
If a connection attempt succeeds, the methods returns and the routine finishes.
|
||||
The same applies when an `ErrCurrentlyDialingOrExistingAddress` error is
|
||||
returned by the dialing method, as it indicates that peer is already connected
|
||||
or that another routine is attempting to (re)connect to it.
|
||||
|
||||
A first set of connection attempts is done at (about) regular intervals.
|
||||
More precisely, between two attempts, the switch waits for a interval of
|
||||
`reconnectInterval`, hard-coded to 5 seconds, plus a random jitter up to
|
||||
`dialRandomizerIntervalMilliseconds`, hard-coded to 3 seconds.
|
||||
At most `reconnectAttempts`, hard-coded to 20, are made using this
|
||||
regular-interval approach.
|
||||
|
||||
A second set of connection attempts is done with exponentially increasing
|
||||
intervals.
|
||||
The base interval `reconnectBackOffBaseSeconds` is hard-coded to 3 seconds,
|
||||
which is also the increasing factor.
|
||||
The exponentially increasing dialing interval is adjusted as well by a random
|
||||
jitter up to `dialRandomizerIntervalMilliseconds`.
|
||||
At most `reconnectBackOffAttempts`, hard-coded to 10, are made using this approach.
|
||||
|
||||
> Note: the first sleep interval, to which a random jitter is applied, is 1,
|
||||
> not `reconnectBackOffBaseSeconds`, as the first exponent is `0`...
|
||||
|
||||
## Accepting peers
|
||||
|
||||
The `acceptRoutine` method is a persistent routine that handles connections
|
||||
accepted by the transport configured for the switch.
|
||||
|
||||
The [`Accept`](./transport.md#accept) method of the configured transport
|
||||
returns a `Peer` with which an inbound connection was established.
|
||||
The switch accepts a new peer if the maximum number of inbound peers was not
|
||||
reached, or if the peer was configured as an _unconditional peer_.
|
||||
The maximum number of inbound peers is determined by the `MaxNumInboundPeers`
|
||||
configuration parameter, whose default value is `40`.
|
||||
|
||||
If accepted, the peer is added to the switch using the [`addPeer`](#add-peer) method.
|
||||
If the switch does not accept the established incoming connection, or if the
|
||||
`addPeer` method returns an error, the switch invokes the transport's
|
||||
[`Cleanup`](./transport.md#cleanup) method to clean any resources associated
|
||||
with the peer.
|
||||
|
||||
The transport's `Accept` method can also return a number of errors.
|
||||
Errors of `ErrRejected` or `ErrFilterTimeout` types are ignored,
|
||||
an `ErrTransportClosed` causes the accepted routine to be interrupted,
|
||||
while other errors cause the routine to panic.
|
||||
|
||||
> TODO: which errors can cause the routine to panic?
|
||||
|
||||
## Add peer
|
||||
|
||||
The `addPeer` method adds a peer to the switch,
|
||||
either after dialing (by `addOutboundPeerWithConfig`, called by `DialPeerWithAddress`)
|
||||
a peer and establishing an outbound connection,
|
||||
or after accepting (`acceptRoutine`) a peer and establishing an inbound connection.
|
||||
|
||||
The first step is to invoke the `filterPeer` method.
|
||||
It checks whether the peer is already in the set of connected peers,
|
||||
and whether any of the configured `peerFilter` methods reject the peer.
|
||||
If the peer is already present or it is rejected by any filter, the `addPeer`
|
||||
method fails and returns an error.
|
||||
|
||||
Then, the new peer is started, added to the set of connected peers, and added
|
||||
to all reactors.
|
||||
More precisely, first the new peer's information is first provided to every
|
||||
reactor (`InitPeer` method).
|
||||
Next, the peer's sending and receiving routines are started, and the peer is
|
||||
added to set of connected peers.
|
||||
These two operations can fail, causing `addPeer` to return an error.
|
||||
Then, in the absence of previous errors, the peer is added to every reactor (`AddPeer` method).
|
||||
|
||||
> Adding the peer to the peer set returns a `ErrSwitchDuplicatePeerID` error
|
||||
> when a peer with the same ID is already presented.
|
||||
>
|
||||
> TODO: Starting a peer could be reduced as starting the MConn with that peer?
|
||||
|
||||
## Stop peer
|
||||
|
||||
There are two methods for stopping a peer, namely disconnecting from it, and
|
||||
removing it from the table of connected peers.
|
||||
|
||||
The `StopPeerForError` method is invoked to stop a peer due to an external
|
||||
error, which is provided to method as a generic "reason".
|
||||
|
||||
The `StopPeerGracefully` method stops a peer in the absence of errors or, more
|
||||
precisely, not providing to the switch any "reason" for that.
|
||||
|
||||
In both cases the `Peer` instance is stopped, the peer is removed from all
|
||||
registered reactors, and finally from the list of connected peers.
|
||||
|
||||
> Issue https://github.com/tendermint/tendermint/issues/3338 is mentioned in
|
||||
> the internal `stopAndRemovePeer` method explaining why removing the peer from
|
||||
> the list of connected peers is the last action taken.
|
||||
|
||||
When there is a "reason" for stopping the peer (`StopPeerForError` method)
|
||||
and the peer is a persistent peer, the method creates a routine to attempt
|
||||
reconnecting to the peer address, using the `reconnectToPeer` method.
|
||||
If the peer is an outbound peer, the peer's address is know, since the switch
|
||||
has dialed the peer.
|
||||
Otherwise, the peer address is retrieved from the `NodeInfo` instance from the
|
||||
connection handshake.
|
||||
|
||||
## Add reactor
|
||||
|
||||
The `AddReactor` method registers a `Reactor` to the switch.
|
||||
|
||||
The reactor is associated to the set of channel ids it employs.
|
||||
Two reactors (in the same node) cannot share the same channel id.
|
||||
|
||||
There is a call back to the reactor, in which the switch passes itself to the
|
||||
reactor.
|
||||
|
||||
## Remove reactor
|
||||
|
||||
The `RemoveReactor` method unregisters a `Reactor` from the switch.
|
||||
|
||||
The reactor is disassociated from the set of channel ids it employs.
|
||||
|
||||
There is a call back to the reactor, in which the switch passes `nil` to the
|
||||
reactor.
|
||||
|
||||
## OnStart
|
||||
|
||||
This is a `BaseService` method.
|
||||
|
||||
All registered reactors are started.
|
||||
|
||||
The switch's `acceptRoutine` is started.
|
||||
|
||||
## OnStop
|
||||
|
||||
This is a `BaseService` method.
|
||||
|
||||
All (connected) peers are stopped and removed from the peer's list using the
|
||||
`stopAndRemovePeer` method.
|
||||
|
||||
All registered reactors are stopped.
|
||||
|
||||
## Broadcast
|
||||
|
||||
This method broadcasts a message on a channel, by sending the message in
|
||||
parallel to all connected peers.
|
||||
|
||||
The method spawns a thread for each connected peer, invoking the `Send` method
|
||||
provided by each `Peer` instance with the provided message and channel ID.
|
||||
The return value (a boolean) of these calls are redirected to a channel that is
|
||||
returned by the method.
|
||||
|
||||
> TODO: detail where this method is invoked:
|
||||
> - By the consensus protocol, in `broadcastNewRoundStepMessage`,
|
||||
> `broadcastNewValidBlockMessage`, and `broadcastHasVoteMessage`
|
||||
> - By the state sync protocol
|
||||
@@ -0,0 +1,222 @@
|
||||
# Transport
|
||||
|
||||
The transport establishes secure and authenticated connections with peers.
|
||||
|
||||
The transport [`Dial`](#dial)s peer addresses to establish outbound connections,
|
||||
and [`Listen`](#listen)s in a configured network address
|
||||
to [`Accept`](#accept) inbound connections from peers.
|
||||
|
||||
The transport establishes raw TCP connections with peers
|
||||
and [upgrade](#connection-upgrade) them into authenticated secret connections.
|
||||
The established secret connection is then wrapped into `Peer` instance, which
|
||||
is returned to the caller, typically the [switch](./switch.md).
|
||||
|
||||
## Dial
|
||||
|
||||
The `Dial` method is used by the switch to establish an outbound connection with a peer.
|
||||
It is a synchronous method, which blocks until a connection is established or an error occurs.
|
||||
The method returns an outbound `Peer` instance wrapping the established connection.
|
||||
|
||||
The transport first dials the provided peer's address to establish a raw TCP connection.
|
||||
The dialing maximum duration is determined by `dialTimeout`, hard-coded to 1 second.
|
||||
The established raw connection is then submitted to a set of [filters](#connection-filtering),
|
||||
which can reject it.
|
||||
If the connection is not rejected, it is recorded in the table of established connections.
|
||||
|
||||
The established raw TCP connection is then [upgraded](#connection-upgrade) into
|
||||
an authenticated secret connection.
|
||||
This procedure should ensure, in particular, that the public key of the remote peer
|
||||
matches the ID of the dialed peer, which is part of peer address provided to this method.
|
||||
In the absence of errors,
|
||||
the established secret connection (`conn.SecretConnection` type)
|
||||
and the information about the peer (`NodeInfo` record) retrieved and verified
|
||||
during the version handshake,
|
||||
are wrapped into an outbound `Peer` instance and returned to the switch.
|
||||
|
||||
## Listen
|
||||
|
||||
The `Listen` method produces a TCP listener instance for the provided network
|
||||
address, and spawns an `acceptPeers` routine to handle the raw connections
|
||||
accepted by the listener.
|
||||
The `NetAddress` method exports the listen address configured for the transport.
|
||||
|
||||
The maximum number of simultaneous incoming connections accepted by the listener
|
||||
is bound to `MaxNumInboundPeer` plus the configured number of unconditional peers,
|
||||
using the `MultiplexTransportMaxIncomingConnections` option,
|
||||
in the node [initialization](https://github.com/tendermint/tendermint/blob/29c5a062d23aaef653f11195db55c45cd9e02715/node/node.go#L563).
|
||||
|
||||
This method is called when a node is [started](https://github.com/tendermint/tendermint/blob/29c5a062d23aaef653f11195db55c45cd9e02715/node/node.go#L972).
|
||||
In case of errors, the `acceptPeers` routine is not started and the error is returned.
|
||||
|
||||
## Accept
|
||||
|
||||
The `Accept` method returns to the switch inbound connections established with a peer.
|
||||
It is a synchronous method, which blocks until a connection is accepted or an error occurs.
|
||||
The method returns an inbound `Peer` instance wrapping the established connection.
|
||||
|
||||
The transport handles incoming connections in the `acceptPeers` persistent routine.
|
||||
This routine is started by the [`Listen`](#listen) method
|
||||
and accepts raw connections from a TCP listener.
|
||||
A new routine is spawned for each accepted connection.
|
||||
The raw connection is submitted to a set of [filters](#connection-filtering),
|
||||
which can reject it.
|
||||
If the connection is not rejected, it is recorded in the table of established connections.
|
||||
|
||||
The established raw TCP connection is then [upgraded](#connection-upgrade) into
|
||||
an authenticated secret connection.
|
||||
The established secret connection (`conn.SecretConnection` type),
|
||||
the information about the peer (`NodeInfo` record) retrieved and verified
|
||||
during the version handshake,
|
||||
as well any error returned in this process are added to a queue of accepted connections.
|
||||
This queue is consumed by the `Accept` method.
|
||||
|
||||
> Handling accepted connection asynchronously was introduced due to this issue:
|
||||
> https://github.com/tendermint/tendermint/issues/2047
|
||||
|
||||
## Connection Filtering
|
||||
|
||||
The `filterConn` method is invoked for every new raw connection established by the transport.
|
||||
Its main goal is avoid the transport to maintain duplicated connections with the same peer.
|
||||
It also runs a set of configured connection filters.
|
||||
|
||||
The transports keeps a table `conns` of established connections.
|
||||
The table maps the remote address returned by a generic connection to a list of
|
||||
IP addresses, to which the connection remote address is resolved.
|
||||
If the remote address of the new connection is already present in the table,
|
||||
the connection is rejected.
|
||||
Otherwise, the connection's remote address is resolved into a list of IPs,
|
||||
which are recorded in the established connections table.
|
||||
|
||||
The connection and the resolved IPs are then passed through a set of connection filters,
|
||||
configured via the `MultiplexTransportConnFilters` transport option.
|
||||
The maximum duration for the filters execution, which is performed in parallel,
|
||||
is determined by `filterTimeout`.
|
||||
Its default value is 5 seconds,
|
||||
which can be changed using the `MultiplexTransportFilterTimeout` transport option.
|
||||
|
||||
If the connection and the resolved remote addresses are not filtered out,
|
||||
the transport registers them into the `conns` table and returns.
|
||||
|
||||
In case of errors, the connection is removed from the table of established
|
||||
connections and closed.
|
||||
|
||||
### Errors
|
||||
|
||||
If the address of the new connection is already present in the `conns` table,
|
||||
an `ErrRejected` error with the `isDuplicate` reason is returned.
|
||||
|
||||
If the IP resolution of the connection's remote address fails,
|
||||
an `AddrError` or `DNSError` error is returned.
|
||||
|
||||
If any of the filters reject the connection,
|
||||
an `ErrRejected` error with the `isRejected` reason is returned.
|
||||
|
||||
If the filters execution times out,
|
||||
an `ErrFilterTimeout` error is returned.
|
||||
|
||||
## Connection Upgrade
|
||||
|
||||
The `upgrade` method is invoked for every new raw connection established by the
|
||||
transport that was not [filtered out](#connection-filtering).
|
||||
It upgrades an established raw TCP connection into a secret authenticated
|
||||
connection, and validates the information provided by the peer.
|
||||
|
||||
This is a complex procedure, that can be summarized by the following three
|
||||
message exchanges between the node and the new peer:
|
||||
|
||||
1. Encryption: the nodes produce ephemeral key pairs and exchange ephemeral
|
||||
public keys, from which are derived: (i) a pair of secret keys used to
|
||||
encrypt the data exchanged between the nodes, and (ii) a challenge message.
|
||||
1. Authentication: the nodes exchange their persistent public keys and a
|
||||
signature of the challenge message produced with the their persistent
|
||||
private keys. This allows validating the peer's persistent public key,
|
||||
which plays the role of node ID.
|
||||
1. Version handshake: nodes exchange and validate each other `NodeInfo` records.
|
||||
This records contain, among other fields, their node IDs, the network/chain
|
||||
ID they are part of, and the list of supported channel IDs.
|
||||
|
||||
Steps (1) and (2) are implemented in the `conn` package.
|
||||
In case of success, they produce the secret connection that is actually used by
|
||||
the node to communicate with the peer.
|
||||
An overview of this procedure, which implements the station-to-station (STS)
|
||||
[protocol][sts-paper] ([PDF][sts-paper-pdf]), can be found [here][peer-sts].
|
||||
The maximum duration for establishing a secret connection with the peer is
|
||||
defined by `handshakeTimeout`, hard-coded to 3 seconds.
|
||||
|
||||
The established secret connection stores the persistent public key of the peer,
|
||||
which has been validated via the challenge authentication of step (2).
|
||||
If the connection being upgraded is an outbound connection, i.e., if the node has
|
||||
dialed the peer, the dialed peer's ID is compared to the peer's persistent public key:
|
||||
if they do not match, the connection is rejected.
|
||||
This verification is not performed in the case of inbound (accepted) connections,
|
||||
as the node does not know a priori the remote node's ID.
|
||||
|
||||
Step (3), the version handshake, is performed by the transport.
|
||||
Its maximum duration is also defined by `handshakeTimeout`, hard-coded to 3 seconds.
|
||||
The version handshake retrieves the `NodeInfo` record of the new peer,
|
||||
which can be rejected for multiple reasons, listed [here][peer-handshake].
|
||||
|
||||
If the connection upgrade succeeds, the method returns the established secret
|
||||
connection, an instance of `conn.SecretConnection` type,
|
||||
and the `NodeInfo` record of the peer.
|
||||
|
||||
In case of errors, the connection is removed from the table of established
|
||||
connections and closed.
|
||||
|
||||
### Errors
|
||||
|
||||
The timeouts for steps (1) and (2), and for step (3), are configured as the
|
||||
deadline for operations on the TCP connection that is being upgraded.
|
||||
If this deadline it is reached, the connection produces an
|
||||
`os.ErrDeadlineExceeded` error, returned by the corresponding step.
|
||||
|
||||
Any error produced when establishing a secret connection with the peer (steps 1 and 2) or
|
||||
during the version handshake (step 3), including timeouts,
|
||||
is encapsulated into an `ErrRejected` error with reason `isAuthFailure` and returned.
|
||||
|
||||
If the upgraded connection is an outbound connection, and the peer ID learned in step (2)
|
||||
does not match the dialed peer's ID,
|
||||
an `ErrRejected` error with reason `isAuthFailure` is returned.
|
||||
|
||||
If the peer's `NodeInfo` record, retrieved in step (3), is invalid,
|
||||
or if reports a node ID that does not match peer ID learned in step (2),
|
||||
an `ErrRejected` error with reason `isAuthFailure` is returned.
|
||||
If it reports a node ID equals to the local node ID,
|
||||
an `ErrRejected` error with reason `isSelf` is returned.
|
||||
If it is not compatible with the local `NodeInfo`,
|
||||
an `ErrRejected` error with reason `isIncompatible` is returned.
|
||||
|
||||
## Close
|
||||
|
||||
The `Close` method closes the TCP listener created by the `Listen` method,
|
||||
and sends a signal for interrupting the `acceptPeers` routine.
|
||||
|
||||
This method is called when a node is [stopped](https://github.com/tendermint/tendermint/blob/46badfabd9d5491c78283a0ecdeb695e21785508/node/node.go#L1019).
|
||||
|
||||
## Cleanup
|
||||
|
||||
The `Cleanup` method receives a `Peer` instance,
|
||||
and removes the connection established with a peer from the table of established connections.
|
||||
It also invokes the `Peer` interface method to close the connection associated with a peer.
|
||||
|
||||
It is invoked when the connection with a peer is closed.
|
||||
|
||||
## Supported channels
|
||||
|
||||
The `AddChannel` method registers a channel in the transport.
|
||||
|
||||
The channel ID is added to the list of supported channel IDs,
|
||||
stored in the local `NodeInfo` record.
|
||||
|
||||
The `NodeInfo` record is exchanged with peers in the version handshake.
|
||||
For this reason, this method is not invoked with a started transport.
|
||||
|
||||
> The only call to this method is performed in the `CustomReactors` constructor
|
||||
> option of a node, i.e., before the node is started.
|
||||
> Note that the default list of supported channel IDs, including the default reactors,
|
||||
> is provided to the transport as its original `NodeInfo` record.
|
||||
|
||||
[peer-sts]: https://github.com/tendermint/tendermint/blob/main/spec/p2p/peer.md#authenticated-encryption-handshake
|
||||
[peer-handshake]:https://github.com/tendermint/tendermint/blob/main/spec/p2p/peer.md#tendermint-version-handshake
|
||||
[sts-paper]: https://link.springer.com/article/10.1007/BF00124891
|
||||
[sts-paper-pdf]: https://github.com/tendermint/tendermint/blob/0.1/docs/sts-final.pdf
|
||||
@@ -0,0 +1,239 @@
|
||||
# Types adopted in the p2p implementation
|
||||
|
||||
This document lists the packages and source files, excluding test units, that
|
||||
implement the p2p layer, and summarizes the main types they implement.
|
||||
Types play the role of classes in Go.
|
||||
|
||||
The reference version for this documentation is the branch
|
||||
[`v0.34.x`](https://github.com/tendermint/tendermint/tree/v0.34.x/p2p).
|
||||
|
||||
State of August 2022.
|
||||
|
||||
## Package `p2p`
|
||||
|
||||
Implementation of the p2p layer of Tendermint.
|
||||
|
||||
### `base_reactor.go`
|
||||
|
||||
`Reactor` interface.
|
||||
|
||||
`BaseReactor` implements `Reactor`.
|
||||
|
||||
**Not documented yet**.
|
||||
|
||||
### `conn_set.go`
|
||||
|
||||
`ConnSet` interface, a "lookup table for connections and their ips".
|
||||
|
||||
Internal type `connSet` implements the `ConnSet` interface.
|
||||
|
||||
Used by the [transport](#transportgo) to store connected peers.
|
||||
|
||||
### `errors.go`
|
||||
|
||||
Defines several error types.
|
||||
|
||||
`ErrRejected` enumerates a number of reason for which a peer was rejected.
|
||||
Mainly produced by the [transport](#transportgo),
|
||||
but also by the [switch](#switchgo).
|
||||
|
||||
`ErrSwitchDuplicatePeerID` is produced by the `PeerSet` used by the [switch](#switchgo).
|
||||
|
||||
`ErrSwitchConnectToSelf` is handled by the [switch](#switchgo),
|
||||
but currently is not produced outside tests.
|
||||
|
||||
`ErrSwitchAuthenticationFailure` is handled by the [PEX reactor](#pex_reactorgo),
|
||||
but currently is not produced outside tests.
|
||||
|
||||
`ErrTransportClosed` is produced by the [transport](#transportgo)
|
||||
and handled by the [switch](#switchgo).
|
||||
|
||||
`ErrNetAddressNoID`, `ErrNetAddressInvalid`, and `ErrNetAddressLookup`
|
||||
are parsing a string to create an instance of `NetAddress`.
|
||||
It can be returned in the setup of the [switch](#switchgo)
|
||||
and of the [PEX reactor](#pex_reactorgo),
|
||||
as well when the [transport](#transportgo) validates a `NodeInfo`, as part of
|
||||
the connection handshake.
|
||||
|
||||
`ErrCurrentlyDialingOrExistingAddress` is produced by the [switch](#switchgo),
|
||||
and handled by the switch and the [PEX reactor](#pex_reactorgo).
|
||||
|
||||
### `fuzz.go`
|
||||
|
||||
For testing purposes.
|
||||
|
||||
`FuzzedConnection` wraps a `net.Conn` and injects random delays.
|
||||
|
||||
### `key.go`
|
||||
|
||||
`NodeKey` is the persistent key of a node, namely its private key.
|
||||
|
||||
The `ID` of a node is a string representing the node's public key.
|
||||
|
||||
### `metrics.go`
|
||||
|
||||
Prometheus `Metrics` exposed by the p2p layer.
|
||||
|
||||
### `netaddress.go`
|
||||
|
||||
Type `NetAddress` contains the `ID` and the network address (IP and port) of a node.
|
||||
|
||||
The API of the [address book](#addrbookgo) receives and returns `NetAddress` instances.
|
||||
|
||||
This source file was adapted from [`btcd`](https://github.com/btcsuite/btcd),
|
||||
a Go implementation of Bitcoin.
|
||||
|
||||
### `node_info.go`
|
||||
|
||||
Interface `NodeInfo` stores the basic information about a node exchanged with a
|
||||
peer during the handshake.
|
||||
|
||||
It is implemented by `DefaultNodeInfo` type.
|
||||
|
||||
The [switch](#switchgo) stores the local `NodeInfo`.
|
||||
|
||||
The `NodeInfo` of connected peers is produced by the
|
||||
[transport](#transportgo) during the handshake, and stored in [`Peer`](#peergo) instances.
|
||||
|
||||
### `peer.go`
|
||||
|
||||
Interface `Peer` represents a connected peer.
|
||||
|
||||
It is implemented by the internal `peer` type.
|
||||
|
||||
The [transport](#transportgo) API methods return `Peer` instances,
|
||||
wrapping established secure connection with peers.
|
||||
|
||||
The [switch](#switchgo) API methods receive `Peer` instances.
|
||||
The switch stores connected peers in a `PeerSet`.
|
||||
|
||||
The [`Reactor`](#base_reactorgo) methods, invoked by the switch, receive `Peer` instances.
|
||||
|
||||
### `peer_set.go`
|
||||
|
||||
Interface `IPeerSet` offers methods to access a table of [`Peer`](#peergo) instances.
|
||||
|
||||
Type `PeerSet` implements a thread-safe table of [`Peer`](#peergo) instances,
|
||||
used by the [switch](#switchgo).
|
||||
|
||||
The switch provides limited access to this table by returing a `IPeerSet`
|
||||
instance, used by the [PEX reactor](#pex_reactorgo).
|
||||
|
||||
### `switch.go`
|
||||
|
||||
Documented in [switch](./switch.md).
|
||||
|
||||
The `Switch` implements the [peer manager](./peer_manager.md) role for inbound peers.
|
||||
|
||||
[`Reactor`](#base_reactorgo)s have access to the `Switch` and may invoke its methods.
|
||||
This includes the [PEX reactor](#pex_reactorgo).
|
||||
|
||||
### `transport.go`
|
||||
|
||||
Documented in [transport](./transport.md).
|
||||
|
||||
The `Transport` interface is implemented by `MultiplexTransport`.
|
||||
|
||||
The [switch](#switchgo) contains a `Transport` and uses it to establish
|
||||
connections with peers.
|
||||
|
||||
### `types.go`
|
||||
|
||||
Aliases for p2p's `conn` package types.
|
||||
|
||||
## Package `p2p.conn`
|
||||
|
||||
Implements the connection between Tendermint nodes,
|
||||
which is encrypted, authenticated, and multiplexed.
|
||||
|
||||
### `connection.go`
|
||||
|
||||
Implements the `MConnection` type and the `Channel` abstraction.
|
||||
|
||||
A `MConnection` multiplexes a generic network connection (`net.Conn`) into
|
||||
multiple independent `Channel`s, used by different [`Reactor`](#base_reactorgo)s.
|
||||
|
||||
A [`Peer`](#peergo) stores the `MConnection` instance used to interact with a
|
||||
peer, which multiplex a [`SecretConnection`](#secret_connectiongo).
|
||||
|
||||
### `conn_go110.go`
|
||||
|
||||
Support for go 1.10.
|
||||
|
||||
### `secret_connection.go`
|
||||
|
||||
Implements the `SecretConnection` type, which is an encrypted authenticated
|
||||
connection built atop a raw network (TCP) connection.
|
||||
|
||||
A [`Peer`](#peergo) stores the `SecretConnection` established by the transport,
|
||||
which is the underlying connection multiplexed by [`MConnection`](#connectiongo).
|
||||
|
||||
As briefly documented in the [transport](./transport.md#Connection-Upgrade),
|
||||
a `SecretConnection` implements the Station-To-Station (STS) protocol.
|
||||
|
||||
The `SecretConnection` type implements the `net.Conn` interface,
|
||||
which is a generic network connection.
|
||||
|
||||
## Package `p2p.mock`
|
||||
|
||||
Mock implementations of [`Peer`](#peergo) and [`Reactor`](#base_reactorgo) interfaces.
|
||||
|
||||
## Package `p2p.mocks`
|
||||
|
||||
Code generated by `mockery`.
|
||||
|
||||
## Package `p2p.pex`
|
||||
|
||||
Implementation of the [PEX reactor](./pex.md).
|
||||
|
||||
### `addrbook.go`
|
||||
|
||||
Documented in [address book](./addressbook.md).
|
||||
|
||||
This source file was adapted from [`btcd`](https://github.com/btcsuite/btcd),
|
||||
a Go implementation of Bitcoin.
|
||||
|
||||
### `errors.go`
|
||||
|
||||
A number of errors produced and handled by the [address book](#addrbookgo).
|
||||
|
||||
`ErrAddrBookNilAddr` is produced by the address book, but handled (logged) by
|
||||
the [PEX reactor](#pex_reactorgo).
|
||||
|
||||
`ErrUnsolicitedList` is produced and handled by the [PEX protocol](#pex_reactorgo).
|
||||
|
||||
### `file.go`
|
||||
|
||||
Implements the [address book](#addrbookgo) persistence.
|
||||
|
||||
### `known_address.go`
|
||||
|
||||
Type `knownAddress` represents an address stored in the [address book](#addrbookgo).
|
||||
|
||||
### `params.go`
|
||||
|
||||
Constants used by the [address book](#addrbookgo).
|
||||
|
||||
### `pex_reactor.go`
|
||||
|
||||
Implementation of the [PEX reactor](./pex.md), which is a [`Reactor`](#base_reactorgo).
|
||||
|
||||
This includes the implementation of the [PEX protocol](./pex-protocol.md)
|
||||
and of the [peer manager](./peer_manager.md) role for outbound peers.
|
||||
|
||||
The PEX reactor also manages an [address book](#addrbookgo) instance.
|
||||
|
||||
## Package `p2p.trust`
|
||||
|
||||
Go documentation of `Metric` type:
|
||||
|
||||
> // Metric - keeps track of peer reliability
|
||||
> // See tendermint/docs/architecture/adr-006-trust-metric.md for details
|
||||
|
||||
Not imported by any other Tendermint source file.
|
||||
|
||||
## Package `p2p.upnp`
|
||||
|
||||
This package implementation was taken from "taipei-torrent".
|
||||
|
||||
It is used by the `probe-upnp` command of the Tendermint binary.
|
||||
Reference in New Issue
Block a user