scylladb

Author	SHA1	Message	Date
Avi Kivity	b858a4669d	cql3: expr: break up expression.hh header Adding a function declaration to expression.hh causes many recompilations. Reduce that by: - moving some restrictions-related definitions to the existing expr/restrictions.hh - moving evaluation related names to a new header expr/evaluate.hh - move utilities to a new header expr/expr-utilities.hh expression.hh contains only expression definitions and the most basic and common helpers, like printing.	2023-06-22 14:21:03 +03:00
Avi Kivity	32b27d6a08	cql3: expr: change evaluation_input vector components to take spans Spans are slightly cleaner, slightly faster (as they avoid an indirection), and allow for replacing some of the arguments with small_vector:s. Closes #14313	2023-06-22 11:28:01 +02:00
Avi Kivity	3a2d8175fb	cql3: update_parameters: use evaluation_inputs compatible row prefetch update_parameters::prefetch_data is used for some list updates (which need a read-before-write to determine the key to update) and for LWT compare-and-swap. Currently they use a custom structure for representing a read row. Switch to the same structure that is used in evaluation_inputs (and in SELECT statement evaluation) to the expression machinery can be reused. The expression representation is irregular (with different fields for the keys and regular/static columns), so we introduce an old_row structure to hold both the clustering key and the regular row values for cas_request. A nice bonus is that we can use get_non_pk_values() to read the data into the format expected by evaluation_inputs, but on the other hand we have to adjust get_prefetched_list() to fix up the type of the returned list (we return it as a map, not a list, so list updates can access the index).	2023-02-12 17:25:41 +02:00
Konstantin Osipov	670b2562a1	lwt: Cassandrda compatibility when incarnating a row for UPDATE When evaluating an LWT condition involving both static and non-static cells, and matching no regular row, the static row must be used UNLESS the IF condition is IF EXISTS/IF NOT EXISTS, in which case special rules apply. Before this fix, Scylla used to assume a row doesn't exist if there is no matching primary key. In Cassandra, if there is a non-empty static row in the partition, a regular row based on the static row' cell values is created in this case, and then this row is used to evaluate the condition. This problem was reported as gh-10081. The reason for Scylla behaviour before the patch was that when implementing LWT I tried to converge Cassandra data model (or lack of thereof) with a relational data model, and assumed a static row is a "shared" portion of a regular row, i.e. a storage level concept intended to save space, and doesn't have independent existence. This was an oversimplification. This patch fixes gh-10081, making Scylla semantics match the one of Cassandra. I will now list other known examples when a static row has an own independent existence as part of a table, for cataloguing purposes. SELECT * from a partition which has a partition key and a static cell set returns 1 row. If later a regular row is added to the partition, the SELECT would still return 1 row, i.e. the static row will disappear, and a regular row will appear instead. Another example showing a static row has an independent existence below: CREATE TABLE t (p int, c int, s int static, PRIMARY KEY(p, c)); INSERT INTO t (p, c) VALUES(1, 1); INSERT INTO t (p, s) VALUES(1, 1) IF NOT EXISTS; In Cassandra (and Scylla), IF NOT EXISTS evaluates to TRUE, even though both the regular row and the partition exist. But the static cells are not set, and the insert only provides a partition key, so the database assumes the insert is operating against a static row. It would be wrong to assume that a static row exists when the partition key exists: INSERT INTO t (p, c, s) VALUES(1, 1, 1) IF NOT EXISTS; [applied] \| p \| c \| s -----------+---+---+------ False \| 1 \| 1 \| null evaluates to False, i.e. the regular row does exist when p and c exist. Issue CREATE TABLE t (p INT, c INT, r INT, s INT static, PRIMARY KEY(p, c)) INSERT INTO t (p, s) VALUES (1, 1); UPDATE t SET s=2, r=1 WHERE p=1 AND c=1 IF s=1 and r=null; - in this case, even though the regular row doesn't exist, the static row does, and should be used for condition evaluation. In other words, IF EXISTS/IF NOT EXISTS have contextual semantics. They apply to the regular row if clustering key is used in the WHERE clause, otherwise they apply to static row. One analogy for static rows is that it is like a static member of C++ or Java class. It's an attribute of the class (assuming class = partition), which is accessible through every object of the class (object = regular row). It is also present if there are no objects of the class, but the class itself exists: i.e. a partition could have no regular rows, but some static cells set, in this case it has a static row. Unlike C++/Java static class members a static row is an optional attribute of the partition. A partition may exist, but the static row may be absent (e.g. no static cell is set). If the static row does exist, all regular rows share its contents, even if they do not exist. A regular row exists when its clustering key is present in the table. A static row exists when at least one static cell is set. Tests are updated because now when no matching row is found for the update we show the value of the static row as the previous value, instead of a non-matching clustering row. Changes in v2: - reworded the commit message - added select tests Closes #10711	2022-06-16 19:23:46 +03:00
Avi Kivity	5937b1fa23	treewide: remove empty comments in top-of-files After `fcb8d040` ("treewide: use Software Package Data Exchange (SPDX) license identifiers"), many dual-licensed files were left with empty comments on top. Remove them to avoid visual noise. Closes #10562	2022-05-13 07:11:58 +02:00
Avi Kivity	fcb8d040e8	treewide: use Software Package Data Exchange (SPDX) license identifiers Instead of lengthy blurbs, switch to single-line, machine-readable standardized (https://spdx.dev) license identifiers. The Linux kernel switched long ago, so there is strong precedent. Three cases are handled: AGPL-only, Apache-only, and dual licensed. For the latter case, I chose (AGPL-3.0-or-later and Apache-2.0), reasoning that our changes are extensive enough to apply our license. The changes we applied mechanically with a script, except to licenses/README.md. Closes #9937	2022-01-18 12:15:18 +01:00
Avi Kivity	a55b434a2b	treewide: extent copyright statements to present day	2021-06-06 19:18:49 +03:00
Michał Chojnowski	3c98806df9	cql3: update_parameters: don't linearize in prefetch_data_builder::add_cell We can deserialize directly from fragmented buffers now.	2020-12-04 09:19:39 +01:00
Wojciech Mitros	45215746fe	increase the maximum size of query results to 2^64 Currently, we cannot select more than 2^32 rows from a table because we are limited by types of variables containing the numbers of rows. This patch changes these types and sets new limits. The new limits take effect while selecting all rows from a table - custom limits of rows in a result stay the same (2^32-1). In classes which are being serialized and used in messaging, in order to be able to process queries originating from older nodes, the top 32 bits of new integers are optional and stay at the end of the class - if they're absent we assume they equal 0. The backward compatibility was tested by querying an older node for a paged selection, using the received paging_state with the same select statement on an upgraded node, and comparing the returned rows with the result generated for the same query by the older node, additionally checking if the paging_state returned by the upgraded node contained new fields with correct values. Also verified if the older node simply ignores the top 32 bits of the remaining rows number when handling a query with a paging_state originating from an upgraded node by generating and sending such a query to an older node and checking the paging_state in the reply(using python driver). Fixes #5101.	2020-08-03 17:32:49 +02:00
Vladimir Davydov	934a87999f	cql: turn prefetch_data::row into struct This will allow us to add helper methods and store extra info in each row. For example, we can add a method for checking if a row has static columns. Also, to build CAS result set, we need to differentiate rows fetched to check conditions from those fetched for reading operations. Using struct as row container will allow us to store this information in each prefetched row.	2019-10-28 21:12:52 +03:00
Konstantin Osipov	a2b629c3a1	lwt: boost update_parameters to serve as a CAS result set In modification_statement/batch_statement, we need to prefetch data to 1) apply list operations 2) evaluate CAS conditions 3) return CAS result set. Boost update_parameters::prefetch_data to serve as a single result set for all of the above. In case of a batch, store multiple rows for multiple clustering keys involved in the batch. Use an ordered set for columns and rows to make sure 3) CAS result set is returned to the client in an ordered manner. Deserialize the primary key and add it to result set rows since it is returned to the client as part of CAS result set. Index columns using ordinal_id - this allows having a single set for all columns and makes columns easy to look up. Remove an extra memcpy to build view objects when looking up a cell by primary key, use partition_key/clustering_key objects for lookup.	2019-10-16 15:56:50 +03:00
Konstantin Osipov	a4ccbece5c	lwt: remove an unnecessary optional around prefetch_data Get rid of an unnecessary optional around update_parameters::prefetch_data. update_parameters won't own prefetch_data in the future anyway, since prefetch_data can be shared among multiple modification statements of a batch, each statement having its own options and hence its own update_parameters instance.	2019-10-16 15:48:25 +03:00
Konstantin Osipov	7a399ebe0d	lwt: move prefetch_data_builder to update_parameters.cc Move prefetch_data_builder class from modification_statement.cc to update_parameters.cc. We're going to share the same builder to build a result set for condition evaluation and to apply updates of batch statements, so we need to share it. No other changes.	2019-10-16 15:48:08 +03:00
Duarte Nunes	05731cb5ad	cql3/lists: Fix multi-cell static list updates in the presence of ckeys This patch fixes a regression introduced in `9e88b60ef5`, which broke the lookup for prefetched values of lists when a clustering key is specified. This is the code that was removed from some list operations: std::experimental::optional<clustering_key> row_key; if (!column.is_static()) { row_key = clustering_key::from_clustering_prefix(*params._schema, prefix); } ... auto&& existing_list = params.get_prefetched_list(m.key().view(), row_key, column); Put it back, in the form of common code in the update_parameters class. Fixes #3703 Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2018-08-20 21:39:37 +01:00
Duarte Nunes	9e88b60ef5	mutation: Set cell using clustering_key_prefix Change the clustering key argument in mutation::set_cell from exploded_clustering_prefix to clustering_key_prefix, which allows for some overall code simplification and fewer copies. This mostly affects the cql3 layer. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-05-04 15:59:50 +02:00
Pekka Enberg	38a54df863	Fix pre-ScyllaDB copyright statements People keep tripping over the old copyrights and copy-pasting them to new files. Search and replace "Cloudius Systems" with "ScyllaDB". Message-Id: <1460013664-25966-1-git-send-email-penberg@scylladb.com>	2016-04-08 08:12:47 +03:00
Tomasz Grabiec	63006e5dd2	query: Serialize collection cells using CQL format We want the format of query results to be eventually defined in the IDL and be independent of the format we use in memory to represent collections. This change is a step in this direction. The change decouples format of collection cells in query results from our in-memory representation. We currently use collection_mutation_view, after the change we will use CQL binary protocol format. We use that because it requires less transformations on the coordinator side. One complication is that some list operations need to retrieve keys used in list cells, not only values. To satisfy this need, new query option was added called "collections_as_maps" which will cause lists and sets to be reinterpreted as maps matching their underlying representation. This allows the coordinator to generate mutations referencing existing items in lists.	2016-02-15 17:05:55 +01:00
Tomasz Grabiec	383296c05b	cql3: Fix handling of lists with static columns List operations and prefetching were not handling static columns correctly. One issue was that prefetching was attaching static column data to row data using ids which might overlap with clustered columns. Another problem was that list operations were always constructing clustering key even if they worked on a static column. For static columns the key would be always empty and lookup would fail. The effect was that list operations which depend on curent state had no effect. Similar problem could be observed on C* 2.1.9, but not on 2.2.3. Fixes #903.	2016-02-15 17:05:55 +01:00
Avi Kivity	79f7431a03	db: change collection_mutation::{one,view} not to use nested classes Nested classes cannot be forward-declared, so change the naming not to use them. Follows atomic_cell{,_view}.	2015-11-13 17:13:07 +02:00
Avi Kivity	d5cf0fb2b1	Add license notices	2015-09-20 10:43:39 +03:00
Tomasz Grabiec	878a740b9d	db: Write query results in serialized form This gives about 30% increase in tps in: build/release/tests/perf/perf_simple_query -c1 --query-single-key This patch switches query result format from a structured one to a serialized one. The problems with structured format are: - high level of indirection (vector of vectors of vectors of blobs), which is not CPU cache friendly - high allocation rate due to fine-grained object structure On replica side, the query results are probably going to be serialized in the transport layer anyway, so this change only subtracts work. There is no processing of the query results on replica other than concatenation in case of range queries. If query results are collected in serialized form from different cores, we can concatenate them without copying by simply appending the fragments into the packet. This optimization is not implemented yet. On coordinator side, the query results would have to be parsed from the transport layer buffers anyway, so this also doesn't add work, but again saves allocations and copying. The CQL server doesn't need complex data structures to process the results, it just goes over it linearly consuming it. This patch provides views, iterators and visitors for consuming query results in serialized form. Currently the iterators assume that the buffer is contiguous but we could easily relax this in future so that we can avoid linearization of data received from seastar sockets. The coordinator side could be optimized even further for CQL queries which do not need processing (eg. select * from cf where ...) we could make the replica send the query results in the format which is expected by the CQL binary protocol client. So in the typical case the coordinator would just pass the data using zero-copy to the client, prepending a header. We do need structure for prefetched rows (needed by list manipulations), and this change adds query result post-processing which converts serialized query result into a structured one, tailored particularly for prefetched rows needs. This change also introduces partition_slice options. In some queries (maybe even in typical ones), we don't need to send partition or clustering keys back to the client, because they are already specified in the query request, and not queried for. The query results hold now keys as optional elements. Also, meta-data like cell timestamp and ttl is now also optional. It is only needed if the query has writetime() or ttl() functions in it, which it typically won't have.	2015-04-15 20:44:50 +02:00

21 Commits