fix(s3/lifecycle): divide cluster budget by active workers, not all capable

gemini pointed out that s3_lifecycle has MaxJobsPerDetection=1
(handler.go:189) — it's a singleton job, only one worker is ever active.
Dividing the cluster_deletes_per_second budget by the count of capable
executors gave the single active worker just 1/N of the configured cap.

Pass adminRuntime.MaxJobsPerDetection through to the decorator. Divisor
is now min(executors, maxJobsPerDetection), clamped to >=1. For
s3_lifecycle (maxJobs=1) the active worker gets the full budget; for a
hypothetical parallel-dispatch job (maxJobs>1) the budget divides
across the running-set.

Tests swap the SharedEvenly case for two pinned scenarios:
  - SingletonJobGetsFullBudget: maxJobs=1 across 4 executors => 100/1
  - SharedEvenlyWhenParallelLimited: maxJobs=4 across 4 executors => 25/worker
  - MaxJobsExceedsExecutors: maxJobs=10 across 4 executors => divisor 4
This commit is contained in:
Chris Lu
2026-05-11 19:05:56 -07:00
parent c51db540cc
commit b85af3483e
3 changed files with 70 additions and 23 deletions
+6 -4
View File
@@ -650,10 +650,12 @@ func (r *Plugin) executeJobWithExecutor(
}
// Apply per-job-type cluster-allocation decoration (e.g. s3_lifecycle
// divides cluster_deletes_per_second by the worker count and ships
// the share via ClusterContext.Metadata). No-op for job types
// without an allocator registered.
clusterContext = r.decorateClusterContextForJob(clusterContext, job.JobType, adminConfigValues)
// divides cluster_deletes_per_second by min(workers, maxJobsPerDetection)
// and ships the share via ClusterContext.Metadata). No-op for job
// types without an allocator registered. MaxJobsPerDetection caps
// the divisor so a singleton job (maxJobs=1) gets the full budget on
// the single active worker, not 1/N of it.
clusterContext = r.decorateClusterContextForJob(clusterContext, job.JobType, adminConfigValues, int(adminRuntime.GetMaxJobsPerDetection()))
completedCh := make(chan *plugin_pb.JobCompleted, 1)
r.pendingExecutionMu.Lock()