Previously we just looped over and loaded each entity one by one. It's
better to try to limit it to fewer queries (ideally one). Fortunately,
Hibernate supports bulk load of entities with both compound and simple
IDs.
This makes tracking whether or not an entity has been persisted in the
database much simpler, since we can use null as a pure sentinel value to
say "this object has not been persisted yet". This makes things like
"updateAll" much easier and simpler to implement efficiently later
because that will rely on knowing what entities are inserts vs merges
When inserting multiple entities, if sorting isn't enabled, then
Hibernate will just add them in the order received. This is suboptimal
for batch query execution. Hibernate flushes the batch as soon as it
"encounters" a new table, so if we have inserts into tables A, B, A, B
in that order, it'll execute each one separately, rather than batching
(A,A), (B,B).
The downside to enabling ordering is that you need to either have
deferred foreign keys (which we have) or Hibernate-style entity
references, e.g. `@OneToOne Entity entity` rather than `VKey<String>
entityId`.
While any future potential migration to Spanner would require that we
remove the deferred foreign keys, this change is orthogonal to that --
we'll need to change the references to direct Hibernate references
anyway which Hibernate will know how to order properly.
This basically controls the number of child entities that will be
batch-loaded when we bulk-load multiple root entities. It will be useful
for situations like RDE where we load a bunch of domains, because it
means we'll batch-load the child entities like nsHosts and gracePeriods.
- Adjusts SLA analysis soak durations to 30 minutes per phase for sandbox and 1 hour per phase for production.
- Updates non-prod automations to auto-advance through canary-1 and canary-5 in crash and sandbox.
- Adds an automation rule to auto-promote releases from crash to sandbox (still need to manually approve the sandbox release).
- Adds documentation in skaffold.yaml clarifying dual deployment.
This uses the already-existing quota manager to acquire quota, if
necessary, before running a flow. If we can't acquire quota, we throw a
command use EPP exception.
This is a version of the quota manager that essentially special-cases
DomainCreateFlow, assuming that's the only flow that we wish to
throttle. If we wish to throttle more flows in the future, we can
re-architect this a bit.
I'm taking these from the Spinnaker pipeline. In prod we need to promote
the artifacts produced to "live" so that, for instance, Dataflow
pipelines get promoted/deployed.
The package was removed, see the errors:
```
Package google-cloud-sdk-app-engine-java is not available, but is referred to by another package.
This may mean that the package is missing, has been obsoleted, or
is only available from another source
However the following packages replace it:
google-cloud-cli-app-engine-java
```
Redo the exported drop list mechanism by replacing the legacy per-TLD
mode in ExportDomainListsAction with a dedicated, once-daily
ExportDropListAction.
Key changes:
1. Reverted ExportDomainListsAction to unconditionally export
single-column active registered domains and deprecated the
INCLUDE_PENDING_DELETE_DATE_FOR_DOMAINS feature flag.
2. Implemented ExportDropListAction at /_dr/task/exportDropList to query
the read replica for upcoming deletions on all open TLDs (where
invoicing is enabled) and output an alphabetically sorted CSV file
(domain_name,tld,deletion_time) to a designated Google Drive folder.
3. Added domainDropListDriveFolderId configuration setting and provider.
4. Registered the action in RequestComponent, routing.txt, and Cloud
Scheduler tasks for production and sandbox.
5. Added comprehensive test coverage in ExportDropListActionTest and
cleaned up legacy test cases in ExportDomainListsActionTest.
BUG=b/553658111
Expose the --expiry_access_period_enabled CLI flag on registrar mutation
commands and drop temporary database-level default constraints for XAP.
Specifically:
- Expose the --expiry_access_period_enabled parameter in
CreateOrUpdateRegistrarCommand and pass it to
Registrar.Builder.setExpiryAccessPeriodEnabled.
- Add expiryAccessPeriodTransitions with default DISABLED to example.yaml.
- Add unit tests for --expiry_access_period_enabled in
CreateRegistrarCommandTest and UpdateRegistrarCommandTest.
- Add Flyway migrations V229 and V230 to drop the temporary database-level
DEFAULT constraints on Tld.expiry_access_period_transitions and
Registrar.expiry_access_period_enabled per db/README.md.
- Regenerate flyway.txt, nomulus.golden.sql, and ER diagrams.
TAG=agy
BUG=http://b/437398822
We're still getting 429 errors after increasing the number of retries to
four. Using 9 means that the maximum number of seconds we wait will be
255 -- 8 wait intervals starting at 1 second means
1+2+4+8+16+32+64+128=255 seconds.
Also log to see if they give us a useful retry-after header
We don't want to have to pass the entire script to Valkey every single
time. Instead, we compute the SHA-1 hash of the script and refer to it
by hash, significantly reducing the amount of bytes we need to send
to the server. If the script is not loaded in Valkey (NOSCRIPT error),
we reload it and retry.
This separates the logic of "how many tokens should a particular
user/group have" from "manage the quota given whatever limits,
contacting Valkey".
This is in preparation for allowing other types of quota management and
throttling besides just on the EPP server (e.g. domain-create
throttling).
We also convert the expirations from seconds to milliseconds (and use a
Duration so that this is masked from users). It's better to have the API
use a full-fledged Duration object rather than an int, and this allows
for finer control over expiration times.
Note that we'll probably want to use a sliding window in the future
instead of a fixed window, but that's a problem for future us.
In PR #3153, Angular dependencies were upgraded to 22.x, which requires Node.js >= 24.15.0 or >= 22.22.3, and build.gradle was updated to use Node 24.19.0. However, the Cloud Build builder container image script release/builder/build.sh was missed and remained on 22.12.0, causing the :console-webapp:buildConsoleWebapp task to fail during weekly deployment builds on Google Cloud Build.
BUG= http://b/547933107
Currently, we process (repoId, revisionId) pairs for DomainHistory and
HostHistory individually -- they may be farmed out to worker nodes in
parallel, but each EppResource uses a separate transaction and a separate read,
which doesn't scale well when there are lots of domains/hosts. So as a
result, we should batch them up so we can load (by default) 500 per
transaction at a time.
We don't want to batch-load the domains/hosts at the same time that we
retrieve the most recent history entry for each type -- this would mean
passing relatively large objects across pipeline steps. Instead, we keep
passing the KV<String, Long> and batch retrievals.
This isn't necessarily much faster (due to having to wait on batching)
but there'll be less load on the DB.
Self-scan D.2 number 5
We're getting some 429 (too many requests) responses from the
SafeBrowsing API which is causing the Spec11 pipeline to fail. Currently
the retrier is set to have a backoff starting at 100ms and only have a
couple retries before failing (the latter is already configurable). For
429s, we want to wait significantly longer. Let's start at one second of
waiting and allow four doublings.
- node to the most recent LTS version
- angular dependencies to 22.x rather than 21.x
- eslint to 10.7 (the old version was quite old)
Deployed to alpha and everything looks good. The Typescript changes
were all done by the migration (nothing manually).
G.1 numbers 8 and 9
We can make the "select the correct most recent history object" query
much more efficient + simple but we'll want to have these indexes so we
can select the most recently modified entry for a given repo ID (at
least, one before a particular watermark).
Change the rotatePrimaryCert field in CreateOrUpdateRegistrarCommand
from object Boolean to primitive boolean. Since the field is initialized
to default false and used as a presence-based switch without tri-state
semantics, primitive boolean accurately reflects its behavior and avoids
unnecessary object wrapper overhead.
BUG= http://b/537308816
For reserved/premium lists:
Use double-check locking so that subsequent calls to get the entire map
of entries don't need to even check the locking object. This makes
things quicker and removes lock-tracking overhead.
For TMCH CA:
we can just remove the synchronization block entirely. Everything inside
of it is either constants (e.g. ROOT_CERTS) or a Guava loading cache
(CRL_CACHE) which takes care of synchronization for us anyway.
Calculates the delay seconds automatically. This value helps ensure that
all EPP requests are enqueued before the scheduled test start time.
Since queue insertion is much slower than dispatch, this is essential to
maintain a stable QPS rate.
Also parallelizes queue insertion using a thread pool. This reduces the
delay for enqueuing the requests.
BUG=http://b/533414332
This is configured to run every 5 minutes. We need to make sure that the
cache doesn't get too out of date, otherwise we'll be serving stale
data. We'll add an alert that fires if SUCCESS or NOT_CONFIGURED hasn't
happened recently.
* Fix image replacement in cd (#3186)
* read sql jobs from ar
* revert release change
* flatten file path for sql jobs
* no source to sql command
* add automation to pipeline
* fix automation
* fix replica seize for backend and console in partial phases
Per ICANN's Expired Registration Recovery Policy, all gTLD registries must
offer a Redemption Grace Period (RGP) of 30 days during which deleted
domains may be restored. Registry reservation lists should not block
domain restore commands during the RGP.
This change removes the reserved list check in DomainRestoreRequestFlow,
reverting the behavior originally added in CL 72341125 (July 2014) that
explicitly disallowed restoring reserved domains. Unit tests have been
updated to confirm restoring reserved domains succeeds for standard
registrar accounts.
BUG=b/539548743
TAG=agy
CONV=d5dff534-f924-4bba-a58e-74091dc5f496
This means we don't have to load all users and filter them out later. In
practice this doesn't matter because the user table is relatively small
(a few hundred) but 1. who knows what can happen in the future? 2. this
makes the code analysis tools happier