mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-07-28 02:55:24 +00:00
* admin/topology: carry the volume server address on DiskInfo The planning DiskInfo exposed only the node id, which can be an opaque label rather than ip:port. Record the address too so callers can resolve the physical machine a disk sits on. * ec.balance: spread a volume's shards across machines, not just nodes Volume servers sharing a host are one fault domain, but the within-rack spread treated them as independent nodes, so one box could end up holding more shards of a volume than EC can afford to lose. Add a machine (host) tier between rack and node: the within-rack pass spreads each volume across machines, and the global load phase no longer re-concentrates a volume onto a machine it already sits on. Host defaults to the node id, so clusters with one server per host are unchanged. * ec placement: prefer machines holding fewer of a volume's shards EC allocation and repair picked the least-loaded node in a rack with no regard for which physical machine it sits on, so a volume's shards could pile onto several servers of one box. Rank candidate nodes by their machine's shard count first, then the node's own. The machine is derived from the volume server address carried on DiskInfo, falling back to the node id, matching how the balancer resolves it. * volume.balance: don't move a replica onto a machine already holding one isGoodMove only rejected a move onto the same data node, so two replicas could land on two volume servers of one box and a single machine failure would lose both. Reject a target whose host already holds another replica of the volume. Best-effort: balancing simply skips and tries the next target. * volume allocation: spread same-rack replicas across machines PickNodesByWeight filled the same-rack replica picks by weight alone, so replicas could co-locate on one box. Prefer candidates on not-yet-used hosts, falling back when too few distinct machines exist. Data-center and rack tiers have no host, so their ordering is unchanged. * ec.balance: harden machine spread against re-concentration and capped machines Two cases where the machine-aware spread could still leave a volume badly placed: - The global load phase could move a shard of a volume onto a machine that already held it, raising that machine's count and undoing the within-rack spread (a 4/4/3/3 layout could become 3/5/3/3, past parity for 10+4). Limit the load-only fallback to same-machine moves, which leave a machine's count unchanged; cross-machine concentration is no longer allowed for load alone. - The within-rack spread chose a destination machine by free slots alone, so if that machine's only nodes were already at the SameRackCount cap it skipped the move instead of trying another machine. Require a machine to have a node that can actually take the shard before selecting it. * reduce comments across the machine-affinity change Trim narration down to the non-obvious why; one terse line where a block was overkill. * ec.balance: gate machine spread on fault-tolerance feasibility Spreading a volume evenly across machines only helps when there are enough that each can stay within EC's parity tolerance (numMachines >= ceil(total/parity)). With fewer -- or wildly unequal -- machines it can't make a machine loss survivable anyway, and forcing it fights capacity: e.g. a cluster of 12 volume servers on one host and 2 on another would have half of every volume crammed onto the 2-server box. So spread across machines only when it's achievable; otherwise fall back to per-node spread and let capacity/global balancing decide. The global load phase applies the same test: it protects a volume's machine spread (no cross-machine move that raises a machine's count past the source's) only where that spread is achievable, so heterogeneous clusters still level by fullness. * ec.balance worker: group servers by host when planning The worker built its planner topology without recording each server's host, so automated ec.balance treated ports on one machine as independent nodes and could concentrate a volume's shards on one physical box. Set the host from the volume server address, matching the shell path. * volume.balance worker: don't move a replica onto a machine holding one The worker compared only node ids, and the replica map dropped the server address, so it could move replicas onto different ports of one machine. Carry the host on ReplicaLocation (from the server address) and reject a target whose host already holds another replica of the volume. Best-effort, matching the shell. * ec.balance: judge machine-spread feasibility by the rack's shards The within-rack and global feasibility checks compared the whole volume's shard count against a rack's machine count, so a rack holding only part of a volume after cross-rack spreading -- e.g. 7 of a 10+4 volume across 2 machines -- was wrongly judged infeasible and fell back to node spread, which could pile 6 shards onto one host, past parity. Gate on the rack's own shard count of the volume instead. * ec.balance: spread a volume's shards across machines by combined count EC recovers from any loss within parity regardless of shard type, so what bounds a machine's exposure is its total shards of the volume, not data and parity separately. Spreading the two independently let each type's remainder land on the same machine -- ceil(d/M)+ceil(p/M) can exceed ceil(total/M), e.g. a 5/3 split where 4/4 was achievable, past parity. Balance the combined count in one pass; disk-level data/parity anti-affinity stays in pickBestDiskOnNode. * ec.balance: don't let the imbalance threshold skip an over-parity machine The within-rack spread gated on relative skew ((max-min)/avg > threshold), so a worker threshold of 0.5 skipped an exactly-50%-skewed layout like 5/4/3 for a 10+4 volume, leaving 5 shards -- past parity -- on one machine. The even cap (ceil(shards/groups)) is the real bound and the move loop already sheds only what exceeds it, so drop the threshold gate from the within-rack phase (machine and node): a balanced rack stays a no-op while any over-cap machine is always fixed. * ec.balance: keep the imbalance threshold for the node fallback Dropping the threshold from the whole within-rack phase made the node fallback too eager: it runs only when machine fault tolerance is unachievable, so it is cosmetic load distribution that should defer to the global utilization phase. Without the gate it would, for a one-server-per-host 6/4 split at threshold 0.5, schedule a count move that worsens utilization balance. Restore the threshold there; machine spreading keeps bypassing it, since that bound is durability, not cosmetic skew.
548 lines
17 KiB
Go
548 lines
17 KiB
Go
package topology
|
|
|
|
import (
|
|
"errors"
|
|
"fmt"
|
|
"math/rand/v2"
|
|
"strings"
|
|
"sync"
|
|
"sync/atomic"
|
|
"time"
|
|
|
|
"github.com/seaweedfs/seaweedfs/weed/glog"
|
|
"github.com/seaweedfs/seaweedfs/weed/stats"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
|
|
"github.com/seaweedfs/seaweedfs/weed/storage/types"
|
|
)
|
|
|
|
type NodeId string
|
|
|
|
// CapacityReservation represents a temporary reservation of capacity
|
|
type CapacityReservation struct {
|
|
reservationId string
|
|
diskType types.DiskType
|
|
count int64
|
|
createdAt time.Time
|
|
}
|
|
|
|
// CapacityReservations manages capacity reservations for a node
|
|
type CapacityReservations struct {
|
|
sync.RWMutex
|
|
reservations map[string]*CapacityReservation
|
|
reservedCounts map[types.DiskType]int64
|
|
}
|
|
|
|
func newCapacityReservations() *CapacityReservations {
|
|
return &CapacityReservations{
|
|
reservations: make(map[string]*CapacityReservation),
|
|
reservedCounts: make(map[types.DiskType]int64),
|
|
}
|
|
}
|
|
|
|
func (cr *CapacityReservations) removeReservation(reservationId string) bool {
|
|
cr.Lock()
|
|
defer cr.Unlock()
|
|
|
|
if reservation, exists := cr.reservations[reservationId]; exists {
|
|
delete(cr.reservations, reservationId)
|
|
cr.decrementCount(reservation.diskType, reservation.count)
|
|
return true
|
|
}
|
|
return false
|
|
}
|
|
|
|
func (cr *CapacityReservations) getReservedCount(diskType types.DiskType) int64 {
|
|
cr.RLock()
|
|
defer cr.RUnlock()
|
|
|
|
return cr.reservedCounts[diskType]
|
|
}
|
|
|
|
// decrementCount is a helper to decrement reserved count and clean up zero entries
|
|
func (cr *CapacityReservations) decrementCount(diskType types.DiskType, count int64) {
|
|
cr.reservedCounts[diskType] -= count
|
|
// Clean up zero counts to prevent map growth
|
|
if cr.reservedCounts[diskType] <= 0 {
|
|
delete(cr.reservedCounts, diskType)
|
|
}
|
|
}
|
|
|
|
// doAddReservation is a helper to add a reservation, assuming the lock is already held
|
|
func (cr *CapacityReservations) doAddReservation(diskType types.DiskType, count int64) string {
|
|
now := time.Now()
|
|
reservationId := fmt.Sprintf("%s-%d-%d-%d", diskType, count, now.UnixNano(), rand.Int64())
|
|
cr.reservations[reservationId] = &CapacityReservation{
|
|
reservationId: reservationId,
|
|
diskType: diskType,
|
|
count: count,
|
|
createdAt: now,
|
|
}
|
|
cr.reservedCounts[diskType] += count
|
|
return reservationId
|
|
}
|
|
|
|
// tryReserveAtomic atomically checks available space and reserves if possible
|
|
func (cr *CapacityReservations) tryReserveAtomic(diskType types.DiskType, count int64, availableSpaceFunc func() int64) (reservationId string, success bool) {
|
|
cr.Lock()
|
|
defer cr.Unlock()
|
|
|
|
// Check available space under lock
|
|
currentReserved := cr.reservedCounts[diskType]
|
|
availableSpace := availableSpaceFunc() - currentReserved
|
|
|
|
if availableSpace >= count {
|
|
// Create and add reservation atomically
|
|
return cr.doAddReservation(diskType, count), true
|
|
}
|
|
|
|
return "", false
|
|
}
|
|
|
|
func (cr *CapacityReservations) cleanExpiredReservations(expirationDuration time.Duration) {
|
|
cr.Lock()
|
|
defer cr.Unlock()
|
|
|
|
now := time.Now()
|
|
for id, reservation := range cr.reservations {
|
|
if now.Sub(reservation.createdAt) > expirationDuration {
|
|
delete(cr.reservations, id)
|
|
cr.decrementCount(reservation.diskType, reservation.count)
|
|
glog.V(1).Infof("Cleaned up expired capacity reservation: %s", id)
|
|
}
|
|
}
|
|
}
|
|
|
|
type Node interface {
|
|
Id() NodeId
|
|
String() string
|
|
AvailableSpaceFor(option *VolumeGrowOption) int64
|
|
ReserveOneVolume(r int64, option *VolumeGrowOption) (*DataNode, error)
|
|
ReserveOneVolumeForReservation(r int64, option *VolumeGrowOption) (*DataNode, error)
|
|
UpAdjustDiskUsageDelta(diskType types.DiskType, diskUsage *DiskUsageCounts)
|
|
UpAdjustMaxVolumeId(vid needle.VolumeId)
|
|
GetDiskUsages() *DiskUsages
|
|
|
|
// Capacity reservation methods for avoiding race conditions
|
|
TryReserveCapacity(diskType types.DiskType, count int64) (reservationId string, success bool)
|
|
ReleaseReservedCapacity(reservationId string)
|
|
AvailableSpaceForReservation(option *VolumeGrowOption) int64
|
|
|
|
GetMaxVolumeId() needle.VolumeId
|
|
SetParent(Node)
|
|
LinkChildNode(node Node)
|
|
UnlinkChildNode(nodeId NodeId)
|
|
CollectDeadNodeAndFullVolumes(freshThreshHold int64, volumeSizeLimit uint64, growThreshold float64)
|
|
|
|
IsDataNode() bool
|
|
IsRack() bool
|
|
IsDataCenter() bool
|
|
IsLocked() bool
|
|
Children() []Node
|
|
Parent() Node
|
|
|
|
GetValue() interface{} //get reference to the topology,dc,rack,datanode
|
|
}
|
|
|
|
type NodeImpl struct {
|
|
diskUsages *DiskUsages
|
|
id NodeId
|
|
parent Node
|
|
sync.RWMutex // lock children
|
|
children map[NodeId]Node
|
|
// maxVolumeId uses atomic ops so UpAdjustMaxVolumeId (called from the
|
|
// volume server heartbeat path) and GetMaxVolumeId (called from the
|
|
// master's assign / warmup checks) can run concurrently without a
|
|
// data race on NodeImpl.
|
|
maxVolumeId atomic.Uint32
|
|
|
|
//for rack, data center, topology
|
|
nodeType string
|
|
value interface{}
|
|
|
|
// capacity reservations to prevent race conditions during volume creation
|
|
capacityReservations *CapacityReservations
|
|
}
|
|
|
|
func (n *NodeImpl) GetDiskUsages() *DiskUsages {
|
|
return n.diskUsages
|
|
}
|
|
|
|
// nodeHost returns the host a node runs on, or "" for non-data-node tiers (data
|
|
// centers, racks).
|
|
func nodeHost(node Node) string {
|
|
if dn, ok := node.(*DataNode); ok {
|
|
return dn.Ip
|
|
}
|
|
return ""
|
|
}
|
|
|
|
// preferDistinctHosts reorders candidates so the first node of each not-yet-used
|
|
// host comes first (preserving weighted order), then the same-host leftovers, so a
|
|
// prefix covers the most distinct machines. usedHost seeds the set. No-op when
|
|
// hosts are empty (non-data-node tiers).
|
|
func preferDistinctHosts(usedHost string, candidates []Node) []Node {
|
|
used := map[string]bool{}
|
|
if usedHost != "" {
|
|
used[usedHost] = true
|
|
}
|
|
distinct := make([]Node, 0, len(candidates))
|
|
dup := make([]Node, 0, len(candidates))
|
|
for _, node := range candidates {
|
|
h := nodeHost(node)
|
|
if h != "" && !used[h] {
|
|
used[h] = true
|
|
distinct = append(distinct, node)
|
|
} else {
|
|
dup = append(dup, node)
|
|
}
|
|
}
|
|
return append(distinct, dup...)
|
|
}
|
|
|
|
// the first node must satisfy filterFirstNodeFn(), the rest nodes must have one free slot
|
|
func (n *NodeImpl) PickNodesByWeight(numberOfNodes int, option *VolumeGrowOption, filterFirstNodeFn func(dn Node) error) (firstNode Node, restNodes []Node, err error) {
|
|
var totalWeights int64
|
|
var errs []string
|
|
n.RLock()
|
|
candidates := make([]Node, 0, len(n.children))
|
|
candidatesWeights := make([]int64, 0, len(n.children))
|
|
//pick nodes which has enough free volumes as candidates, and use free volumes number as node weight.
|
|
for _, node := range n.children {
|
|
if node.AvailableSpaceFor(option) <= 0 {
|
|
continue
|
|
}
|
|
totalWeights += node.AvailableSpaceFor(option)
|
|
candidates = append(candidates, node)
|
|
candidatesWeights = append(candidatesWeights, node.AvailableSpaceFor(option))
|
|
}
|
|
n.RUnlock()
|
|
if len(candidates) < numberOfNodes {
|
|
glog.V(0).Infoln(n.Id(), "failed to pick", numberOfNodes, "from ", len(candidates), "node candidates")
|
|
return nil, nil, errors.New("Not enough data nodes found!")
|
|
}
|
|
|
|
//pick nodes randomly by weights, the node picked earlier has higher final weights
|
|
sortedCandidates := make([]Node, 0, len(candidates))
|
|
for i := 0; i < len(candidates); i++ {
|
|
// Break if no more weights available to prevent panic in rand.Int64N
|
|
if totalWeights <= 0 {
|
|
break
|
|
}
|
|
weightsInterval := rand.Int64N(totalWeights)
|
|
lastWeights := int64(0)
|
|
for k, weights := range candidatesWeights {
|
|
if (weightsInterval >= lastWeights) && (weightsInterval < lastWeights+weights) {
|
|
sortedCandidates = append(sortedCandidates, candidates[k])
|
|
candidatesWeights[k] = 0
|
|
totalWeights -= weights
|
|
break
|
|
}
|
|
lastWeights += weights
|
|
}
|
|
}
|
|
|
|
restNodes = make([]Node, 0, numberOfNodes-1)
|
|
ret := false
|
|
n.RLock()
|
|
for k, node := range sortedCandidates {
|
|
if err := filterFirstNodeFn(node); err == nil {
|
|
firstNode = node
|
|
// Fill the rest preferring not-yet-used hosts, so replicas spread across
|
|
// machines; falls back to same-host when too few. No-op for dc/rack tiers
|
|
// (empty host), which keep the weighted order.
|
|
pool := make([]Node, 0, len(sortedCandidates)-1)
|
|
pool = append(pool, sortedCandidates[:k]...)
|
|
pool = append(pool, sortedCandidates[k+1:]...)
|
|
pool = preferDistinctHosts(nodeHost(firstNode), pool)
|
|
if len(pool) > numberOfNodes-1 {
|
|
pool = pool[:numberOfNodes-1]
|
|
}
|
|
restNodes = pool
|
|
ret = true
|
|
break
|
|
} else {
|
|
errs = append(errs, string(node.Id())+":"+err.Error())
|
|
}
|
|
}
|
|
n.RUnlock()
|
|
if !ret {
|
|
return nil, nil, errors.New("No matching data node found! \n" + strings.Join(errs, "\n"))
|
|
}
|
|
return
|
|
}
|
|
|
|
func (n *NodeImpl) IsDataNode() bool {
|
|
return n.nodeType == "DataNode"
|
|
}
|
|
|
|
func (n *NodeImpl) IsRack() bool {
|
|
return n.nodeType == "Rack"
|
|
}
|
|
|
|
func (n *NodeImpl) IsDataCenter() bool {
|
|
return n.nodeType == "DataCenter"
|
|
}
|
|
|
|
func (n *NodeImpl) IsLocked() (isTryLock bool) {
|
|
if isTryLock = n.TryRLock(); isTryLock {
|
|
n.RUnlock()
|
|
}
|
|
return !isTryLock
|
|
}
|
|
|
|
func (n *NodeImpl) String() string {
|
|
if n.parent != nil {
|
|
return n.parent.String() + ":" + string(n.id)
|
|
}
|
|
return string(n.id)
|
|
}
|
|
|
|
func (n *NodeImpl) Id() NodeId {
|
|
return n.id
|
|
}
|
|
|
|
func (n *NodeImpl) getOrCreateDisk(diskType types.DiskType) *DiskUsageCounts {
|
|
return n.diskUsages.getOrCreateDisk(diskType)
|
|
}
|
|
|
|
func (n *NodeImpl) AvailableSpaceFor(option *VolumeGrowOption) int64 {
|
|
t := n.getOrCreateDisk(option.DiskType)
|
|
freeVolumeSlotCount := atomic.LoadInt64(&t.maxVolumeCount) + atomic.LoadInt64(&t.remoteVolumeCount) - atomic.LoadInt64(&t.volumeCount)
|
|
freeVolumeSlotCount -= ecShardSlots(atomic.LoadInt64(&t.ecShardCount))
|
|
return freeVolumeSlotCount
|
|
}
|
|
|
|
// AvailableSpaceForReservation returns available space considering existing reservations
|
|
func (n *NodeImpl) AvailableSpaceForReservation(option *VolumeGrowOption) int64 {
|
|
baseAvailable := n.AvailableSpaceFor(option)
|
|
reservedCount := n.capacityReservations.getReservedCount(option.DiskType)
|
|
return baseAvailable - reservedCount
|
|
}
|
|
|
|
// TryReserveCapacity attempts to atomically reserve capacity for volume creation
|
|
func (n *NodeImpl) TryReserveCapacity(diskType types.DiskType, count int64) (reservationId string, success bool) {
|
|
const reservationTimeout = 5 * time.Minute // TODO: make this configurable
|
|
|
|
// Clean up any expired reservations first
|
|
n.capacityReservations.cleanExpiredReservations(reservationTimeout)
|
|
|
|
// Atomically check and reserve space
|
|
option := &VolumeGrowOption{DiskType: diskType}
|
|
reservationId, success = n.capacityReservations.tryReserveAtomic(diskType, count, func() int64 {
|
|
return n.AvailableSpaceFor(option)
|
|
})
|
|
|
|
if success {
|
|
glog.V(1).Infof("Reserved %d capacity for diskType %s on node %s: %s", count, diskType, n.Id(), reservationId)
|
|
}
|
|
|
|
return reservationId, success
|
|
}
|
|
|
|
// ReleaseReservedCapacity releases a previously reserved capacity
|
|
func (n *NodeImpl) ReleaseReservedCapacity(reservationId string) {
|
|
if n.capacityReservations.removeReservation(reservationId) {
|
|
glog.V(1).Infof("Released capacity reservation on node %s: %s", n.Id(), reservationId)
|
|
} else {
|
|
glog.V(1).Infof("Attempted to release non-existent reservation on node %s: %s", n.Id(), reservationId)
|
|
}
|
|
}
|
|
func (n *NodeImpl) SetParent(node Node) {
|
|
n.parent = node
|
|
}
|
|
|
|
func (n *NodeImpl) Children() (ret []Node) {
|
|
n.RLock()
|
|
defer n.RUnlock()
|
|
for _, c := range n.children {
|
|
ret = append(ret, c)
|
|
}
|
|
return ret
|
|
}
|
|
|
|
func (n *NodeImpl) Parent() Node {
|
|
return n.parent
|
|
}
|
|
|
|
func (n *NodeImpl) GetValue() interface{} {
|
|
return n.value
|
|
}
|
|
|
|
func (n *NodeImpl) ReserveOneVolume(r int64, option *VolumeGrowOption) (assignedNode *DataNode, err error) {
|
|
return n.reserveOneVolumeInternal(r, option, false)
|
|
}
|
|
|
|
// ReserveOneVolumeForReservation selects a node using reservation-aware capacity checks
|
|
func (n *NodeImpl) ReserveOneVolumeForReservation(r int64, option *VolumeGrowOption) (assignedNode *DataNode, err error) {
|
|
return n.reserveOneVolumeInternal(r, option, true)
|
|
}
|
|
|
|
func (n *NodeImpl) reserveOneVolumeInternal(r int64, option *VolumeGrowOption, useReservations bool) (assignedNode *DataNode, err error) {
|
|
n.RLock()
|
|
defer n.RUnlock()
|
|
for _, node := range n.children {
|
|
var freeSpace int64
|
|
if useReservations {
|
|
freeSpace = node.AvailableSpaceForReservation(option)
|
|
} else {
|
|
freeSpace = node.AvailableSpaceFor(option)
|
|
}
|
|
// fmt.Println("r =", r, ", node =", node, ", freeSpace =", freeSpace)
|
|
if freeSpace <= 0 {
|
|
continue
|
|
}
|
|
if r >= freeSpace {
|
|
r -= freeSpace
|
|
} else {
|
|
var hasSpace bool
|
|
if useReservations {
|
|
hasSpace = node.IsDataNode() && node.AvailableSpaceForReservation(option) > 0
|
|
} else {
|
|
hasSpace = node.IsDataNode() && node.AvailableSpaceFor(option) > 0
|
|
}
|
|
if hasSpace {
|
|
// fmt.Println("vid =", vid, " assigned to node =", node, ", freeSpace =", node.FreeSpace())
|
|
dn := node.(*DataNode)
|
|
if dn.IsTerminating {
|
|
continue
|
|
}
|
|
return dn, nil
|
|
}
|
|
if useReservations {
|
|
assignedNode, err = node.ReserveOneVolumeForReservation(r, option)
|
|
} else {
|
|
assignedNode, err = node.ReserveOneVolume(r, option)
|
|
}
|
|
if err == nil {
|
|
return
|
|
}
|
|
}
|
|
}
|
|
return nil, errors.New("No free volume slot found!")
|
|
}
|
|
|
|
func (n *NodeImpl) UpAdjustDiskUsageDelta(diskType types.DiskType, diskUsage *DiskUsageCounts) { //can be negative
|
|
existingDisk := n.getOrCreateDisk(diskType)
|
|
existingDisk.addDiskUsageCounts(diskUsage)
|
|
if n.parent != nil {
|
|
n.parent.UpAdjustDiskUsageDelta(diskType, diskUsage)
|
|
}
|
|
}
|
|
func (n *NodeImpl) UpAdjustMaxVolumeId(vid needle.VolumeId) {
|
|
target := uint32(vid)
|
|
for {
|
|
current := n.maxVolumeId.Load()
|
|
if current >= target {
|
|
return
|
|
}
|
|
if n.maxVolumeId.CompareAndSwap(current, target) {
|
|
break
|
|
}
|
|
}
|
|
if n.parent != nil {
|
|
n.parent.UpAdjustMaxVolumeId(vid)
|
|
}
|
|
}
|
|
func (n *NodeImpl) GetMaxVolumeId() needle.VolumeId {
|
|
return needle.VolumeId(n.maxVolumeId.Load())
|
|
}
|
|
|
|
func (n *NodeImpl) LinkChildNode(node Node) {
|
|
n.Lock()
|
|
defer n.Unlock()
|
|
n.doLinkChildNode(node)
|
|
}
|
|
|
|
func (n *NodeImpl) doLinkChildNode(node Node) {
|
|
if n.children[node.Id()] == nil {
|
|
n.children[node.Id()] = node
|
|
for dt, du := range node.GetDiskUsages().usages {
|
|
n.UpAdjustDiskUsageDelta(dt, du)
|
|
}
|
|
n.UpAdjustMaxVolumeId(node.GetMaxVolumeId())
|
|
node.SetParent(n)
|
|
// Maintain the topology's address index so Ping admission and other
|
|
// callers can resolve a data node from its address in O(1).
|
|
if dn, ok := node.GetValue().(*DataNode); ok {
|
|
if topo := n.GetTopology(); topo != nil {
|
|
topo.registerDataNodeAddress(dn)
|
|
}
|
|
}
|
|
glog.V(0).Infoln(n, "adds child", node.Id())
|
|
}
|
|
}
|
|
|
|
func (n *NodeImpl) UnlinkChildNode(nodeId NodeId) {
|
|
n.Lock()
|
|
defer n.Unlock()
|
|
node := n.children[nodeId]
|
|
if node != nil {
|
|
// Drop the topology address index before clearing the parent pointer
|
|
// so GetTopology() can still walk up to the root.
|
|
if dn, ok := node.GetValue().(*DataNode); ok {
|
|
if topo := n.GetTopology(); topo != nil {
|
|
topo.unregisterDataNodeAddress(dn.ServerAddress(), dn)
|
|
}
|
|
}
|
|
node.SetParent(nil)
|
|
delete(n.children, node.Id())
|
|
for dt, du := range node.GetDiskUsages().negative().usages {
|
|
n.UpAdjustDiskUsageDelta(dt, du)
|
|
}
|
|
glog.V(0).Infoln(n, "removes", node.Id())
|
|
}
|
|
}
|
|
|
|
func (n *NodeImpl) CollectDeadNodeAndFullVolumes(freshThreshHoldUnixTime int64, volumeSizeLimit uint64, growThreshold float64) {
|
|
if n.IsRack() {
|
|
for _, c := range n.Children() {
|
|
dn := c.(*DataNode) //can not cast n to DataNode
|
|
for _, v := range dn.GetVolumes() {
|
|
topo := n.GetTopology()
|
|
diskType := types.ToDiskType(v.DiskType)
|
|
vl := topo.GetVolumeLayout(v.Collection, v.ReplicaPlacement, v.Ttl, diskType)
|
|
|
|
if v.Size >= volumeSizeLimit {
|
|
vl.accessLock.RLock()
|
|
vacuumTime, ok := vl.vacuumedVolumes[v.Id]
|
|
vl.accessLock.RUnlock()
|
|
|
|
// If a volume has been vacuumed in the past 20 seconds, we do not check whether it has reached full capacity.
|
|
// After 20s(grpc timeout), theoretically all the heartbeats of the volume server have reached the master,
|
|
// the volume size should be correct, not the size before the vacuum.
|
|
if !ok || time.Now().Add(-20*time.Second).After(vacuumTime) {
|
|
//fmt.Println("volume",v.Id,"size",v.Size,">",volumeSizeLimit)
|
|
topo.chanFullVolumes <- v
|
|
}
|
|
} else if float64(v.Size) > float64(volumeSizeLimit)*growThreshold {
|
|
topo.chanCrowdedVolumes <- v
|
|
}
|
|
copyCount := v.ReplicaPlacement.GetCopyCount()
|
|
if copyCount > 1 {
|
|
if copyCount > len(topo.Lookup(v.Collection, v.Id)) {
|
|
stats.MasterReplicaPlacementMismatch.WithLabelValues(v.Collection, v.Id.String()).Set(1)
|
|
} else {
|
|
stats.MasterReplicaPlacementMismatch.WithLabelValues(v.Collection, v.Id.String()).Set(0)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
} else {
|
|
for _, c := range n.Children() {
|
|
c.CollectDeadNodeAndFullVolumes(freshThreshHoldUnixTime, volumeSizeLimit, growThreshold)
|
|
}
|
|
}
|
|
}
|
|
|
|
func (n *NodeImpl) GetTopology() *Topology {
|
|
var p Node = n
|
|
for p.Parent() != nil {
|
|
p = p.Parent()
|
|
}
|
|
// A detached subtree (no Topology root in scope) must not panic; the
|
|
// callers above check the returned value for nil and skip the
|
|
// address-index maintenance in that case.
|
|
topo, _ := p.GetValue().(*Topology)
|
|
return topo
|
|
}
|