Commit Graph
22 Commits
Author SHA1 Message Date
Eric EntzelandBen McClelland 3610eddf40 fix: add explicit RDMA build constraints
Gate native RDMA, cuObject, and cuobjclient implementations behind the
rdma build tag while keeping fallback stubs available for standard builds.
Preserve the separate cuobjclient_host configuration, clarify unsupported
platform errors, and update Makefile RDMA targets to pass the required tags
and disable VCS stamping.

Co-authored-by: Ben McClelland <ben.mcclelland@versity.com>
2026-09-15 11:03:51 -07:00
Jihyeon Gim ae2a3a6d55 rdma: acquire session admission credit atomically
Checking the publication backlog and taking the admission
credit were separate steps, so concurrent registrations could
each observe the same headroom and overshoot the session quota
together. Both now share one critical section, and a concurrent
test pins the behavior: sixteen registrations against a limit
of eight with one record pending admit exactly seven.

An admission refusal now publishes the same SlowDown error the
wire response carries, so operational accounting matches what
the client saw, and unregister releases the credit an
unfinalized registration was holding so the admission budget
cannot leak.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 60bec2c00d rdma: gate session admission on audit publication backlog
The native side releases its session quota when it fires the
teardown notification, before the audit record lands in a sink,
so session turnover can queue more unpublished records than the
live-session limit allows. Hold an admission credit per session
from registration until its record is published, and refuse new
sessions while the backlog of unpublished session records reaches
the native session quota: the refusal rolls the prepare back,
still publishes the request-level audit record, and answers
SlowDown so the client retries. A stalled sink now turns into
latency instead of unbounded memory.

Count dropped request records under the publication mutex so the
shutdown drop-count report cannot miss an increment racing it.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 51e63e2636 rdma: cap the session-less publication backlog
Bound the records a stalled sink can accumulate from requests
that never opened a session (failed authentications): beyond
4096 queued, dispatchOrDrop drops the record and counts it, and
shutdown reports the drop count once. Session publications stay
uncapped - each session publishes exactly once and the session
table has a hard limit, so their backlog is structurally bounded.

Cancel the metrics child context on constructor failure so a
malformed publisher endpoint does not leak the derived context
onto the parent.
2026-09-09 13:16:52 +09:00
Jihyeon Gim c9668b40b7 rdma: keep sinks outside the publication lock and self-close metrics
Detach queued work under the publication mutex and execute the
sinks after releasing it: the overflow servicing and the shutdown
drain now swap out the pending list inside the critical section
and publish outside, so a slow sink delays its own records but
never blocks a dispatcher or a native terminal callback.

Move the worker-stopped transition under the same mutex. The
worker marks itself stopped before emptying the channel and the
overflow list, and dispatchers test that flag rather than the
done channel inside the critical section, closing the window
where an append could land between the sweep completing and the
deferred channel close only to be stranded.

Give the metrics manager its own cancellation. The forwarder
terminates through a child context the manager derives and
cancels in Close, so a standalone use of the API shuts down
without relying on an external context being canceled.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 64a6ed0c37 rdma: synchronize publication acceptance with the drain boundary
Service the overflow list during normal operation: the worker
publishes and clears it after every channel job, so a burst that
exceeds the channel buffer drains as soon as the sink recovers
instead of accumulating until shutdown.

Make dispatch and the drain sweep share one critical section. An
append either lands before the sweep and is drained, or runs
after the worker exited and publishes inline; the check-then-send
window that could strand a record between the two is closed.
Request publications use the same boundary with drop semantics:
once the drain finished, no owner can guarantee the sinks are
still open, so the record is dropped rather than published.

Never close the metrics datapoint channel. The forwarder exits
through the canceled context after draining the buffer, the
closed flag turns late producers into drops, and no send can race
a closure. Implement the new Manager method in the test mock.
2026-09-09 13:16:52 +09:00
Jihyeon Gim e9a2a2f98b rdma: never run operational sinks on callback threads
Rework the publication handoff so a native callback can never
execute a sink: the channel buffer absorbs the common case, and a
full buffer appends to an overflow list that the worker drains
after the channel instead of falling back to inline publication.
Queued records accumulate across successive sessions and failed
authentications consume no session at all, so capacity accounting
cannot bound the backlog; only removing the fallback closes the
stall. Drop the now-unused session-limit accessor.

Move signature verification back outside the admission barrier:
IAM lookups carry no cancellation, so holding the barrier across
verification let one stalled lookup defer RC shutdown
indefinitely. The handlers enforce admission themselves, and a
failure publication checks the drain state before dispatching, so
it cannot land after the sinks close.

Synchronize metrics producers with Close: the manager now marks
itself closed before closing the datapoint channel, and a
producer that still races the closure recovers instead of
panicking on a send over a closed channel.

Exercise real stale tokens in the reservation generation test:
the original token attempts both release and publish after a
newer claim took over.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 55c82f5e4e rdma: size publications to session capacity and preserve committed PUT facts
Size the publication queue from the configured session limit instead
of a fixed depth with an overflow semaphore: each session publishes
exactly one terminal record, so the queue can never fill and native
callbacks hand off without waiting. Remove the semaphore fallback.

Admit verification failures under the shutdown barrier so an
authentication publication cannot land after the sinks close. Give
the metrics manager a lifetime independent of the gateway context so
shutdown drain updates are counted, and pass the captured bucket
explicitly so RC datapoints appear in bucket-filtered metrics.

Track reservations by claim generation: release and publication
validate the generation, so a stale claim cannot consume a newer
owner's record. Install the reservation cleanup defer immediately
after acquisition so a panic during authorization cannot orphan it.

Preserve committed PUT facts independent of the native finalizer:
when the backend object exists, record the commit, keep the
committed byte count on the error publication, and still emit the
object-created event.
2026-09-09 13:16:52 +09:00
Jihyeon Gim d3e45fdb41 rdma: track terminal events held off by an active reservation
A teardown callback that arrived while its record was reserved
was dropped: if the READY that held the reservation then rolled
its claim back, the session had lost both publication paths - no
request owner and no callback owner - and the record stayed
forever, unpublished. The record now stashes the terminal event;
releasing the reservation after a rollback consumes the stash and
publishes the expiry, since the native session is gone and no
second callback will arrive.

A duplicate READY that lost the reservation race still proceeded
to claim the native transfer, so the transfer could complete with
no publication owner. The READY handler now refuses to claim when
the reservation is not granted, answering as a duplicate claim.
Publication also validates ownership against the record's emitter,
so a stale reservation cannot publish over or consume the current
owner's record.

The shutdown drain raced its producers: a callback could enqueue
a publication after the drain checked the queue but before the
worker exited, stranding the record behind a stopped worker. The
drain now runs after the RC service close, which quiesces the
native reaper before returning, so no producer can enqueue behind
the drain; ordering replaces locking.

Error paths that published directly bypassed the single-shot
guard, letting the deferred panic safety net attempt a second
publication of the same record. All outcome publications in the
READY completion flow now go through the guarded path.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 96be4d1407 rdma: bound publication backpressure and close the reservation window
Overflow publications ran one goroutine per job: a stalled sink
with a full queue grew them without bound (a probe reached a
thousand blocked calls). Overflow now runs inline under a bounded
semaphore, so at most a fixed number of callers wait and every
record still publishes.

The publication worker outlived the operational sinks: gateway
shutdown closed the RC service and then the sinks while
publications were still queued, losing terminal records (reached
"file already closed" with the real file logger). The shutdown
wrapper now drains the publication queue before closing the RC
service, and the route handler exposes the drain; late terminals
after the drain publish inline instead of queueing behind a
stopped worker.

The publication reservation happened after the transfer claim
returned: in the window between the claim and the reservation, a
concurrent READY's re-authorization denial consumed the unreserved
record, so the claimant's successful transfer lost its publication
to the denial. The READY handler now reserves before the claim and
before re-authorization; a rolled-back claim (wire failure, peer
busy) releases the reservation instead of publishing, so the
session keeps its record for the next claimant or the reaper.

The tracker test now joins the worker through the shutdown drain
and asserts per-session outcomes and byte counts instead of an
aggregate count that a pending fifth record could satisfy.
2026-09-09 13:16:52 +09:00
Jihyeon Gim d30ace72ad rdma: keep reaper callbacks off sink threads
The teardown callback ran audit, metrics, and event sinks inline on
whatever thread fired it, which is the native reaper thread: one
blocking sink (a synchronous file write on a stalled filesystem)
would stall reaping for every other session. Publications now hand
off to a dedicated worker through a bounded queue; a full queue
falls back to a detached goroutine, so the callback never waits on
a sink and no record is dropped.

A denial issued while a completion owner holds the reservation
(concurrent READY re-authorization failure) no longer steals the
publication: reserved records are invisible to the denial path,
which previously produced a denial record plus the owner's success
record for one transfer.

The tracker tests now drive a recording audit sink and assert the
published record count across expiry, reserved, denied-while-
reserved, and consume-or-noop paths: one record per session.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 97238d80ad rdma: reserve session publication before completion calls
A native completion call fires the teardown callback before it
returns, so an outcome recorded after the call is too late: every
successful transfer published as an expiry, and PUT failures and
panics lost their real outcome to the callback's placeholder.

The READY handler now reserves the session record before invoking
any completion call. A reserved record is invisible to the
callback, and the handler publishes the real outcome exactly once
after the result is known. A deferred safety net publishes on
panic unwinds before the unwind finalizer retires the session.
Successful PUTs forward the backend-assigned ETag and version into
the object-created event, and the audit record carries the
transferred byte count as the object size.

A READY whose re-authorization fails now publishes the denial
itself before canceling, instead of letting the cancel publish an
expiry.

The native side keeps the terminal reason on every failure exit:
a verify failure no longer overwrites the wire-failure outcome,
and QP transition or re-arm failures record the wire failure they
return. The unclaimed-teardown classification follows the wire
mapping, so every transfer-level failure the client would see as
502 publishes as the same 502 instead of a diverging per-cause
code.

The metrics bucket tag is documented as absent on synthesized
publications: route parameters come from route matching, which a
synthesized request never runs; the audit log derives the bucket
from the path and stays accurate.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 23a7121bb8 rdma: make the teardown callback the single RC publication owner
Round-3 review of the ownership model found that moving publication
between the request path and the callback by hand leaves edges where
a record is published twice or lost. The publication model is now
structural instead: request paths only record outcomes, and the
native teardown callback - which the ABI fires exactly once per
destroyed session, after every completion call - is the single
publisher.

The READY handler records its outcome (success with the reported
byte count, or the mapped failure) and lets the callback publish.
A deferred recorder covers panic unwinds, so every path through
the handler leaves a final outcome behind. A failed PREPARE
finalization publishes through a consume-or-noop helper: when the
finalizing call already reaped the session the entry is gone and
the helper is a no-op, otherwise it publishes the failure itself.

Unclaimed teardowns no longer all read as expiries: the record is
classified from the native outcome, so a transfer that died on the
wire, failed verification, or timed out carries its own code and
status.

The synthesized publication context runs on an immutable app, so
string accessors copy instead of exposing the pooled request buffer
to the asynchronously serializing event senders. Captured account
and region strings are cloned for the same reason. The nil-error
publication no longer asserts on the S3 error interface, and the
panic marker is a proper internal S3 error.

Unit tests cover the table semantics: expiry publication,
first-record-wins, consume-or-noop failure publication, unregister,
unknown sessions, error normalization, status mapping, and the
outcome classification.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 279f5e7d2f rdma: harden RC outcome publication ownership and record fidelity
Review of the publication path found that ownership could change
hands at the wrong moment and that records could disagree with both
the wire response and the underlying operation.

Successful transfers lost their completion record: the native
completion calls fire the teardown callback synchronously, so the
callback claimed the publication first and logged every completed
transfer as an expiry, and committed PUTs produced no object-created
events. The READY handler now reserves the publication before
invoking any completion call; a reserved record is invisible to the
callback, and the handler publishes the real outcome exactly once.

The same race existed at creation: the PREPARE finalizer can reap an
expired session and fire the callback before the session is
registered, leaving an orphan entry whose only notification already
happened. Registration now runs before the finalizing call, a
notification that arrives first is parked and consumed by the
registration, and a failed finalization drops the entry.

Retained records referenced the request's pooled header buffer, so
a later request could rewrite a tracked session's bucket and key;
captured strings are cloned now. The synthesized publication path
follows the same rule for the event senders, which serialize
asynchronously.

Operational sinks classify plain errors as 500 on their own, so a
resource-limit rejection logged 500 while the client saw 429. The
publication renders non-S3 errors through the route error mapping
before the record reaches the sinks, and the expiry record carries
a dedicated SessionExpired code instead of a generic one. Malformed
PUT headers now preserve the operation in the record.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 39aae5435c rdma: close publication gaps for RC pre-session failures
Review of the operational publication found four gaps where the
records disagreed with the S3 surface or were missing entirely.

Authorization-failure records lost the requester because the
pre-session publisher did not carry the authenticated account; the
account now flows into the record, so the audit trail names who was
denied.

Signature failures and malformed PREPARE headers ended the request
before any publication point. The auth adapter now publishes
authentication failures through the route handler, and header
validation failures publish with whatever object identity the
headers still carried, matching the S3 surface where every denied
request still logs.

RC requests carried no region: the custom routes run before the
middleware that stores the region local, so event records lacked
awsRegion and audit host headers read s3..amazonaws.com. The auth
adapter sets the region for every verified RC request.

Operational records reported generic 500 statuses for protocol
errors that the wire answers with a specific status (a resource
limit rejection logged 500 while the client saw 429). Status
mapping now reuses the route error mapping, so the recorded status
always equals the wire status.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 49b1c6f6bc rdma: publish RC transfer outcomes into the gateway operational services
Two-phase RC transfers were invisible to the access log, metrics,
and bucket notifications: every outcome record on the S3 surface is
driven by the fiber request context, and the RC wire requests carry
the transfer session rather than the object.

Publish one record per session by tracking the operational context
from PREPARE through completion. A tracker table registers each
successfully prepared session with its account, region, object, and
start time; whichever path confirms the final outcome first (READY
completion, READY failure, or the reaper teardown callback) claims
the entry and publishes exactly once. Sessions that end before
PREPARE succeeds publish a request record directly.

The record is emitted through a synthetic fiber context carrying
the object path and the captured locals, so the existing logger,
metrics manager, and event sender produce the same schema as the
S3 surface without any interface change. Bucket notifications fire
for completed PUTs.

The gateway creates the operational services inside RunVersityGW,
after the RC routes exist. Add an OnServicesReady callback to the
embedded gateway config and wire it in vgwrdma to hand the services
to the RC routes and install the teardown callback.
2026-09-09 13:16:52 +09:00
Jihyeon Gim 6eb574dc1a rdma: expose live RC sessions on the admin server
Add a SessionsSnapshot view over the new C ABI entry point and
serve it from the admin server as GET /rc-sessions. The admin server
gains a WithAdminRoute option so an embedding binary can register
extra admin routes that run with the same signature verification and
admin checks as the built-in endpoints; vgwrdma registers the
snapshot there when the RC feature is enabled. The route replies
with the usual XML error surface so unsigned or non-admin requests
get a 403 rather than a generic 500. Stub builds return a not
supported error, keeping the build matrix unchanged.
2026-09-07 16:14:53 +09:00
Jihyeon Gim 1596531003 rdma: keep RC route XML fidelity
Classify the platform-stub answer before the internal-error
logging decision, so the expected 501 no longer logs as an
internal 500 while debugging production servers.

Serialize the full S3 error XML body instead of the base error
alone: per-type diagnostics such as the access key and the
string-to-sign survive the route boundary. A regression test
wraps a signature failure with both diagnostic fields and
asserts the response keeps the status, the code, and both
fields.

Assert the response body identifiers equal the request-ID
headers, pinning the two views of the same response.
2026-09-04 19:44:45 +09:00
Jihyeon Gim 158ccdfe58 rdma: test RC route errors through the production server
Cover the route error boundary with the real S3 server: the
shared serializer keeps status and body for wrapped S3 errors,
raw fiber errors stay 500, and the platform stub answers 501.
The stub-answer classifier moves next to the shared marker type
in the same commit so every build answers 501 at the point the
test first runs, and the general CI workflow builds the
session-server archive before go test, which the Linux link of
this package now requires.
2026-09-04 19:44:45 +09:00
Jihyeon Gim 50f4482173 rdma: map RC transport errors to protocol codes
Classify the rejected-argument, short-transfer, and oversized-
value failures of the session server as bad requests at the
route boundary, answering the closest S3-style protocol error
instead of a generic internal failure.
2026-09-04 19:44:45 +09:00
Jihyeon Gim 1a8d4c9c97 rdma: serialize RC route errors at the route boundary
The RC control routes returned fiber.Error values for 400, 404,
409, 429, 502, and 503 outcomes, but the production S3 error
handler converts ordinary fiber errors into a generic 500
response, so clients observed InternalError for every protocol
outcome. The routes now send the final status and S3-style XML
body themselves through a shared terminal serializer.

S3-aware errors from authentication, authorization, and the
object backend keep their status and code. Session-server
failures map to protocol error codes: InvalidRdmaRequest,
NoSuchRdmaSession, RdmaSessionConflict, RdmaResourceLimit,
RdmaTransferFailed, and RdmaServiceUnavailable. Owner mismatch
answers the same 404 as an unknown session so a session id is
never disclosed across principals. The platform stub keeps its
501 answer and uses the same response shape.
2026-09-04 19:44:45 +09:00
535cc9d521 feat: add the hipobj-rc-v2 control routes to the vgwrdma gateway
* rdma: add the hipobj-rc-v2 control routes to the vgwrdma gateway

Mount the three control routes (prepare, ready, cancel) on the
S3 port behind the standard SigV4 middleware. The routes own
authentication-adjacent policy the C server cannot see: the
middleware wrapper yields to the handler on success, READY and
CANCEL re-read the account through the IAM cache bypass so
mid-flow deletions and credential rotations take effect
immediately, and every object access re-authorizes against the
decoded bucket and key.

The READY handler implements the session ownership contract:
the completion-reference finalizer installs only after the
transfer claim succeeds, the PUT path hands the reference to
the put view exactly at the borrow point, and the FINAL reply
carries the stored object's metadata. Backend I/O runs under a
context merged with the RC service context so shutdown unblocks
in-flight handlers, with a bounded pool for the fresh IAM
lookups.

vgwrdma starts the session server alongside the gateway when an
RDMA interface is configured, tears it down on exit, and shuts
the IAM service down on any startup failure. embedgw learns the
readonly flag for the object access checks the routes share.

Signed-off-by: Jihyeon Gim <potatogim@potatogim.net>

* rdma: add the missing stub handlers for non-Linux builds

The non-Linux rcroutes stub exposed only Register while the vgwrdma
gateway registers the prepare/ready/cancel handlers directly, so
cross-compiling cmd/vgwrdma failed with undefined methods. Add the
three stub handlers answering 501 Not Implemented and let Register
reuse them, matching the Linux Handler API surface.

* auth: drop the duplicated GetUserAccountFresh definition

The rebase onto main (which already carries GetUserAccountFresh from
the iam-cache-fresh change) kept both copies of the method, breaking
the build with a redeclaration error. Remove the second copy so the
method is defined once.

* rdma: address the review findings on the control route wiring

Drop the unused Handler.Register from both build variants: the
gateway mounts the three control routes through s3api.WithRoute so
the SigV4 verifier wrapper (rcAuth) runs in front of each handler,
and nothing else calls Register.

Clear iamOwned only when RunVersityGW returns nil. It shuts the IAM
service down itself at the end of its shutdown sequence, but its
early failure paths return before reaching that point, so the
deferred shutdown must keep covering those errors.

Remove the unused rcserver.SessionInfo parameter from sizeOf; the
transferred byte count comes from the READY response alone.

* rdma: keep transient IAM failures retryable in the fresh revalidation

The fresh account revalidation turned every GetUserAccountFresh
error into 403, which reports transient backend failures (LDAP
timeouts, network errors) as a revoked account and leaves the
client no room to retry. Only a confirmed missing account
(auth.ErrNoSuchUser) means that; answer anything else with 503 so
clients can retry the request.

* rdma: make the IAM shutdown exactly-once and keep gateway errors visible

The gateway and RunVersityGW share the IAM service, and which side
shut it down could not be told from the return value: runtime
failures return after RunVersityGW already shut the service down,
while early setup failures return before any shutdown happens. The
iamOwned flag therefore either shut the service down twice or leaked
it depending on the error, and the error itself was dropped.

Wrap the service so Shutdown runs exactly once no matter which side
calls it, keep the deferred shutdown for every early failure path,
and return the gateway error again. The wrapper re-exposes the
optional interfaces (fresh account reads, signing keys, policy
evaluation, fixed bucket ownership) so feature detection through the
IAM service keeps working.

* rdma: reuse the SigV4 account for RC control requests

READY and CANCEL are independently authenticated SigV4 requests.
Use the account resolved by the normal SigV4 path instead of
bypassing the IAM cache a second time. This aligns RC revocation
latency with other signed S3 requests and removes the extra
backend IAM lookup, its concurrency cap, and the RC-specific IAM
error mapping. The session owner check and the READY target and
operation authorization are unchanged.

* rdma: reword the READY reauthorization comment

The comment implied a revocation inside the session window always
takes effect at READY, but the account used here is the one SigV4
resolved, which may be a cached entry. State what the check does
without claiming account-cache freshness.

* rdma: preserve IAM cache behavior and standalone region

---------

Signed-off-by: Jihyeon Gim <potatogim@potatogim.net>
Co-authored-by: Ben McClelland <ben.mcclelland@versity.com>
2026-09-02 12:35:53 -07:00