scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-06-03 21:47:10 +00:00

Author	SHA1	Message	Date
Piotr Jastrzebski	76d7c761d1	schema: Stop using deprecated constructor This is another boring patch. One of schema constructors has been deprecated for many years now but was used in several places anyway. Usage of this constructor could lead to data corruption when using MX sstables because this constructor does not set schema version. MX reading/writing code depends on schema version. This patch replaces all the places the deprecated constructor is used with schema_builder equivalent. The schema_builder sets the schema version correctly. Fixes #8507 Test: unit(dev) Signed-off-by: Piotr Jastrzebski <piotr@scylladb.com> Message-Id: <4beabc8c942ebf2c1f9b09cfab7668777ce5b384.1622357125.git.piotr@scylladb.com>	2021-05-30 11:58:27 +03:00
Wojciech Mitros	725c6aac81	test/perf: close test_env to pass an assert in sstables_manager destructor When destroying an perf_sstable_test_env, an assert in sstables_manager destructor fails, because it hasn't been closed. Fix by removing all references to sstables from perf_sstable_test_env, and then closing the test_env(as well as the sstables_manager) Fixes #8736 Signed-off-by: Wojciech Mitros <wojciech.mitros@scylladb.com> Closes #8737	2021-05-27 17:41:17 +03:00
Pavel Emelyanov	d2442a1bb3	tests: Ditch storage_service_for_tests The purpose of the class in question is to start sharded storage service to make its global instance alive. I don't know when exactly it happened but no code that instantiates this wrapper really needs the global storage service. Ref: #2795 tests: unit(dev), perf_sstable(dev) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Message-Id: <20210526170454.15795-1-xemul@scylladb.com>	2021-05-27 14:39:13 +03:00
Pavel Solodovnikov	0663aa6ca1	service_level_controller: remove extraneous `service/storage_service.hh` include Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2021-05-20 02:18:41 +03:00
Pavel Solodovnikov	fff7ef1fc2	treewide: reduce boost headers usage in scylla header files `dev-headers` target is also ensured to build successfully. Signed-off-by: Pavel Solodovnikov <pa.solodovnikov@scylladb.com>	2021-05-20 01:33:18 +03:00
Piotr Sarna	6c6ccda8a0	perf: add alternator frontend to perf_simple_query The perf_simple_query tool is extended with another protocol aside from CQL - alternator. The alternative (pun intended) benchmark can be executed by using the `--alternator X` parameter, where X specifies one of the alternator's mandatory write isolation options: - "forbid_rmw" - forbids RMW (read-modify-write) requests - "unsafe" - never uses LWT (lightweight transactions), even for RMW - "always_use_lwt" - uses LWT even for non-RMW requests - "only_rmw_uses_lwt" - that one's rather self-explanatory Alternator cooperates with existing --write and --delete parameters. Aside from being able to check for improvements/regressions in the alternator module, it's also possible to check how different isolation levels influence the number of allocations and overall performance, or to compare alternator against CQL. $ ./build/release/test/perf/perf_simple_query_g --smp 1 \ --write --alternator only_rmw_uses_lwt --default-log-level error random-seed=1235000092 Started alternator executor 10873.76 tps (202.9 allocs/op, 12.4 tasks/op, 369921 insns/op) 11096.09 tps (202.7 allocs/op, 12.1 tasks/op, 374792 insns/op) 11100.09 tps (203.0 allocs/op, 12.1 tasks/op, 376469 insns/op) 11068.98 tps (203.1 allocs/op, 12.1 tasks/op, 377132 insns/op) 11081.24 tps (203.2 allocs/op, 12.1 tasks/op, 377290 insns/op) median 11081.24 tps (203.2 allocs/op, 12.1 tasks/op, 377290 insns/op) median absolute deviation: 14.85 maximum: 11100.09 minimum: 10873.76 $ ./build/release/test/perf/perf_simple_query_g --smp 1 \ --random-seed 1235000092 --write --alternator always_use_lwt \ --default-log-level error random-seed=1235000092 Started alternator executor 3605.35 tps (877.4 allocs/op, 174.6 tasks/op, 986666 insns/op) 3555.71 tps (890.0 allocs/op, 174.4 tasks/op, 1006945 insns/op) 3530.20 tps (899.7 allocs/op, 174.1 tasks/op, 1021908 insns/op) 3437.65 tps (908.2 allocs/op, 174.6 tasks/op, 1033992 insns/op) 3409.88 tps (913.2 allocs/op, 174.4 tasks/op, 1041240 insns/op) median 3530.20 tps (899.7 allocs/op, 174.1 tasks/op, 1021908 insns/op) median absolute deviation: 75.15 maximum: 3605.35 minimum: 3409.88	2021-05-18 15:10:31 +02:00
Botond Dénes	dca808dd51	perf/perf_simple_query: add --enable-cache option Allowing for testing performance with/out cache. Signed-off-by: Botond Dénes <bdenes@scylladb.com> Message-Id: <20210517045402.16153-1-bdenes@scylladb.com>	2021-05-17 14:06:18 +02:00
Avi Kivity	8d6e575f59	perf_fast_forward: report instructions per fragment Use a hardware counter to report instructions per fragment. Results vary from ~4k insns/f when reading sequentially to more than 1M insns/f. Instructions per fragment can be a more stable metric than frags/sec. It would probably be even more stable with a fake file implementation that works in-memory to eliminate seastar polling instruction variation. Closes #8660	2021-05-17 11:33:24 +02:00
Benny Halevy	f4cfa530cc	perf: enable instructions_retired_counter only once per executor::run Enabling it for each run_worker call will invoke ioctl PERF_EVENT_IOC_ENABLE in parallel to other workers running and this may skew the results. Test: perf_simple_query Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20210514130542.301168-1-bhalevy@scylladb.com>	2021-05-16 12:13:27 +03:00
Tomasz Grabiec	a9dd7a295d	tests: perf_row_cache_reads: Add scenario for lots of rows covered by a range tombstone Reproduces #8626. Output: test_scan_with_range_delete_over_rows Populating with rows Rows: 702710 Scanning... read: 540.007324 [ms], preemption: {count: 2356, 99%: 1.131752 [ms], max: 1.148589 [ms]}, cache: 251/252 [MB] read: 651.942688 [ms], preemption: {count: 1176, 99%: 1.131752 [ms], max: 1.009652 [ms]}, cache: 251/252 [MB]	2021-05-12 11:58:36 +02:00
Avi Kivity	2b252ef9b7	test: perf: tidy up executor_stats snapshot computation Now that executor_stats_snapshot() is a member function, we can move the capture of _count into invocations into it, capturing all the stats in one place.	2021-04-28 19:02:35 +03:00
Avi Kivity	863b49af03	test: perf: report instructions retired per operations Instructions retired per op is a much more stable than time per op (inverse throughput) since it isn't much affected by changes in CPU frequencey or other load on the test system (it's still somewhat affected since a slower system will run more reactor polls per op). It's also less indicative of real performance, since it's possible for fewer inststructions to execute in more time than more instructions, but that isn't an issue for comparative tests). This allows incremental changes to the code base to be compared with more confidence.	2021-04-28 18:46:55 +03:00
Avi Kivity	0bc98caf3e	test: perf: add RAII wrapper around Linux perf_event_open() Make it easy to embed in other classes. A helper function is provided for the instructions retired counter.	2021-04-28 18:41:02 +03:00
Avi Kivity	498e6b9a64	test: perf: make executor_stats_snapshot() a member function of executor I'd like to add an instructions counter which isn't accessible via a global, so make the snapshot function a member. Out of respect to #1, define functions for getting the number of allocations and tasks processed, as they need heavy header files.	2021-04-28 18:38:35 +03:00
Avi Kivity	a43e896396	test: perf: don't truncate allocation/req and tasks/req report I used {:.0} to truncate to integer, but apparently that resulted in only one significant digit in the report, so 93.1 was reported as 90. Use the {:5.1f} to avoid truncation, and even get an extra digit (we can have fractional tasks/op due to batching). Current result is 93.1 allocs/op, 20.1 tasks/op (which suggests batch size of around 10). Closes #8550	2021-04-28 12:50:13 +02:00
Benny Halevy	391f942b2a	perf: everywhere: close flat_mutation_reader when done Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2021-04-25 11:35:07 +03:00
Benny Halevy	844bc40060	everywhere: use with_closeable to close flat_mutation_reader `with_closeable` simplifies scoped use of flat_mutation_reader, making sure to always close the reader after use. Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2021-04-25 11:16:10 +03:00
Avi Kivity	daeddda7cc	treewide: remove inclusions of storage_proxy.hh from headers storage_proxy.hh is huge and includes many headers itself, so remove its inclusions from headers and re-add smaller headers where needed (and storage_proxy.hh itself in source files that need it). Ref #1.	2021-04-20 21:23:00 +03:00
Benny Halevy	9c89702fb2	perf_simple_query: use tests::random::get_int for reproducible results Support for random-seed was added in `4ad06c7eeb` but the program still uses std::rand() to draw random keys. Use tests::random::get_int instead so we can get reprodicible sequence of keys given a particular random-seed. Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20210418104455.82086-1-bhalevy@scylladb.com>	2021-04-19 17:22:53 +02:00
Tomasz Grabiec	320f6bf220	Merge 'test: perf: perf_simple_query: collect allocation and task statistics' from Avi Kivity Calculate and display the number of memory allocations and tasks executed per operation. Sample results (--smp 1): 180022.46 tps (90 allocs/op, 20 tasks/op) 178963.44 tps (90 allocs/op, 20 tasks/op) 178702.41 tps (90 allocs/op, 20 tasks/op) 177679.74 tps (90 allocs/op, 20 tasks/op) 179539.36 tps (90 allocs/op, 20 tasks/op) median 178963.44 tps (90 allocs/op, 20 tasks/op) median absolute deviation: 575.92 maximum: 180022.46 minimum: 177679.74 This allows less noisy tracking of how some changes impact performance. Closes #8425 * github.com:scylladb/scylla: test: perf: perf_simple_query: collect allocation and task statistics perf: deinline some functions in perf.hh	2021-04-14 13:16:00 +02:00
Avi Kivity	202c631dee	test: perf: perf_simple_query: collect allocation and task statistics Calculate and display the number of memory allocations and tasks executed per operation. Sample results (--smp 1): 180022.46 tps (90 allocs/op, 20 tasks/op) 178963.44 tps (90 allocs/op, 20 tasks/op) 178702.41 tps (90 allocs/op, 20 tasks/op) 177679.74 tps (90 allocs/op, 20 tasks/op) 179539.36 tps (90 allocs/op, 20 tasks/op) median 178963.44 tps (90 allocs/op, 20 tasks/op) median absolute deviation: 575.92 maximum: 180022.46 minimum: 177679.74 This allows less noisy tracking of how some changes impact performance.	2021-04-07 17:54:48 +03:00
Avi Kivity	3a90df39c5	perf: deinline some functions in perf.hh Those functions were defined in a header, but not marked inline. This made including the header from two source files impossible, as the linker would complain about duplicate symbols. Rather than making them inline, put them in a new source file perf.cc as they don't need to be inline.	2021-04-07 17:51:58 +03:00
Avi Kivity	29a674cd94	test: perf: perf_fast_forward: report allocation rate and tasks These are more stable than cpu consumed across runs, and impact performance directly. Closes #8422	2021-04-07 15:41:43 +02:00
Pavel Emelyanov	8bbe2eae5e	btree: Convert comparator to <=> It turned out that all the users of btree can already be converted to use safer std::strong_ordering. The only meaningful change here is the btree code itself -- no more ints there. tests: unit(dev) Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Message-Id: <20210330153648.27049-1-xemul@scylladb.com>	2021-04-01 12:56:08 +03:00
Piotr Jastrzebski	57c7964d6c	config: ignore enable_sstables_mc_format flag Don't allow users to disable MC sstables format any more. We would like to retire some old cluster features that has been around for years. Namely MC_SSTABLE and UNBOUNDED_RANGE_TOMBSTONES. To do this we first have to make sure that all existing clusters have them enabled. It is impossible to know that unless we stop supporting enable_sstables_mc_format flag. Test: unit(dev) Refs #8352 Signed-off-by: Piotr Jastrzebski <piotr@scylladb.com> Closes #8360	2021-03-31 12:23:59 +03:00
Piotr Jastrzebski	3aa7bee5e3	perf: Add benchmarks for large partitions in perf_mutation_readers. Signed-off-by: Piotr Jastrzebski <piotr@scylladb.com>	2021-03-29 09:48:11 +02:00
Avi Kivity	5f4bf18387	Revert "Merge 'sstables: add versioning to the sstable_set ' from Wojciech Mitros" This reverts commit `31909515b3`, reversing changes made to `ef97adc72a`. It shows many serious regressions in dtest. Fixes #8197.	2021-03-02 13:21:22 +02:00
Avi Kivity	31909515b3	Merge 'sstables: add versioning to the sstable_set ' from Wojciech Mitros Currently, the sstable_set in a table is copied before every change to allow accessing the unchanged version by existing sstable readers. This patch changes the sstable_set to a structure that keeps all its versions that are referenced somewhere and provides a way of getting a reference to an immutable version of the set. Each sstable in the set is associated with the versions it is alive in, and is removed when all such versions don't have references anymore. To avoid copying, the object holding all sstables in the set version is changed to a new structure, sstable_list, which was previously an alias for std::unordered_set<shared_sstable>, and which implements most of the methods of an unordered_set, but its iterator uses the actual set with all sstables from all referenced versions and iterates over those sstables that belong to the captured version. The methods that modify the sets contents give strong exception guarantee by trying to insert new sstables to its containers, and erasing them in the case of an caught exception. To release shared_sstables as soon as possible (i.e. when all references to versions that contain them die), each time a version is removed, all sstables that were referenced exclusively by this version are erased. We are able to find these sstables efficiently by storing, for each version, all sstables that were added and erased in it, and, when a version is removed, merging it with the next one. When a version that adds an sstable gets merged with a version that removes it, this sstable is erased. Fixes #2622 Signed-off-by: Wojciech Mitros wojciech.mitros@scylladb.com Closes #8111 * github.com:scylladb/scylla: sstables: add test for checking the latency of updating the sstable_set in a table sstables: move column_family_test class from test/boost to test/lib sstables: use fast copying of the sstable_set instead of rebuilding it sstables: replace the sstable_set with a versioned structure sstables: remove potential ub sstables: make sstable_set constructor less error-prone	2021-03-01 14:16:36 +02:00
Avi Kivity	d980f550d1	Merge 'row_cache: Make fill_buffer() preemptable when cursor leads with dummy rows' from Tomasz Grabiec fill_buffer() will keep scanning until _lower_bound_changed is true, even if preemption is signaled, so that the reader makes forward progress. Before the patch, we did not update _lower_bound on touching a dummy entry. The read will not respect preemption until we hit a non-dummy row. If there is a lot of dummy rows, that can cause reactor stalls. Fix that by updating _lower_bound on dummy entries as well. Refs #8153. Tested with perf_row_cache_reads: ``` $ build/release/test/perf/perf_row_cache_reads -c1 -m200M Rows in cache: 0 Populating with dummy rows Rows in cache: 373929 Scanning read: 183.658966 [ms], preemption: {count: 848, 99%: 0.545791 [ms], max: 0.519343 [ms]}, cache: 99/100 [MB] read: 120.951515 [ms], preemption: {count: 257, 99%: 0.545791 [ms], max: 0.518795 [ms]}, cache: 99/100 [MB] ``` Notice that max preemption latency is low in the second "read:" line. Closes #8167 * github.com:scylladb/scylla: row_cache: Make fill_buffer() preemptable when cursor leads with dummy rows tests: perf: Introduce perf_row_cache_reads row_cache: Add metric for dummy row hits	2021-02-28 21:00:20 +02:00
Tomasz Grabiec	52e411df36	tests: perf: Introduce perf_row_cache_reads Tests performance of various read patterns from the row cache. Example: $ build/release/test/perf/perf_row_cache_reads_g -c1 -m200M Filling memtable Rows in cache: 0 Populating with dummy rows Rows in cache: 373929 Scanning read: 156.288986 [ms], preemption: {count: 702, 99%: 0.545791 [ms], max: 0.537537 [ms]}, cache: 99/100 [MB] read: 106.480766 [ms], preemption: {count: 6, 99%: 0.006866 [ms], max: 106.496168 [ms]}, cache: 99/100 [MB]	2021-02-26 01:20:38 +01:00
Pavel Emelyanov	9baf1226dc	test/memory_footpring: Print radix tree node sizes After switching cells storage onto compact radix tree it becomes useful to know the tree nodes' sizes. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 20:41:09 +03:00
Pavel Emelyanov	aa85bc790b	test: Add tests for radix tree Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 20:27:00 +03:00
Pavel Emelyanov	0ad361b380	test/perf_collection: Add callback to check the speed of clone In some places scylla clones collections of objects, so it's sometimes needed to measure the speed of this operation. This patch adds a placeholder for it, but no implementations for any supported collections. It will be added soon for radix tree. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 17:46:37 +03:00
Pavel Emelyanov	767253fe24	test/perf_mutation: Add option to run with more than 1 columns The --column-count makes the test generate schema with the given numbers of columns and make mutation maker fill random column with the value on each iteration. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 17:45:42 +03:00
Pavel Emelyanov	fc84ab3418	test/perf_mutation: Prepare to have several regular columns Teach the schema builder and test itself to work on more than one regular column, but for now only use 1, as before. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 17:44:34 +03:00
Pavel Emelyanov	21adff2a41	test/perf_mutation: Use builder to build schema The test will be taught to use more than one regular column, so switch to builder in advance. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-15 17:44:06 +03:00
Wojciech Mitros	17634d141b	sstables: add test for checking the latency of updating the sstable_set in a table Added a test which measures the time it takes to replace sstables in a table's sstable_set, using the leveled compaction strategy. Signed-off-by: Wojciech Mitros <wojciech.mitros@scylladb.com>	2021-02-11 11:02:55 +01:00
Avi Kivity	37b41d7764	test: add missing #include <fstream> std::ofstream is used, but there is no direct include for it. This fails the build with libstdc++ 11. Closes #8050	2021-02-09 14:45:20 +02:00
Tomasz Grabiec	2e3f6a9622	tests: perf_fast_forward: Print outpout directory Message-Id: <20210203180053.230627-1-tgrabiec@scylladb.com>	2021-02-04 10:39:41 +02:00
Tomasz Grabiec	e0ceb454c0	tests: perf_fast_forward: Print error hints to stdout They point to lines printed to stdout, so should be aligned with them. Message-Id: <20210203180016.230547-1-tgrabiec@scylladb.com>	2021-02-04 10:39:41 +02:00
Pavel Emelyanov	a92eb2f7a9	perf-test : Print B-tree sizes After the switch from BST to B-tree the memory foorprint includes inner/leaf nodes from the B-tree, so it's useful to know their sizes too. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-02 09:30:30 +03:00
Pavel Emelyanov	2f7c03d84c	utils: Intrusive B-tree (with tests) The design of the tree goes from the row-cache needs, which are 1. Insert/Remove do not invalidate iterators 2. Elements are LSA-manageable 3. Low key overhead 4. External tri-comparator 5. As little actions on insert/remove as possible With the above the design is Two types of nodes -- inner and leaf. Both types keep pointer on parent nodes and N pointers on keys (not keys themselves). Two differences: inner nodes have array of pointers on kids, leaf nodes keep pointer on the tree (to update left- and rightmost tree pointers on node move). Nodes do not keep pointers/references on trees, thus we have O(1) move of any object, but O(logN) to get the tree size. Fortunately, with big keys-per-node value this won't result in too many steps. In turn, the tree has 3 pointers -- root, left- and rightmost leaves. The latter is for constant-time begin() and end(). Keys are managed by user with the help of embeddable member_hook instance, which is 1 pointer in size. The code was copied from the B+ tree one, then heavily reworked, the internal algorythms turned out to differ quite significantly. For the sake of mutation_partition::apply_monotonically(), which needs to move an element from one tree into another, there's a key_grabber helping wrapper that allows doing this move respecting the exception-safety requirement. As measured by the perf_collections test the B-tree with 8 keys is faster, than the std::set, but slower than the B+tree: vs set vs b+tree fill: +13% -6% find: +23% -35% Another neat thing is that 1-key insertion-removal is ~40% faster than for BST (the same number of allocations, but the key object is smaller, less pointers to set-up and less instructions to execute when linking node with root). v4: - equip insertion methods with on_alloc_point() calls to catch potential exception guarantees violations eariler - add unlink_leftmost_without_rebalance. The method is borrowed from boost intrusive set, and is added to kill two birds -- provide it, as it turns out to be popular, and use a bit faster step-by-step tree destruction than plain begin+erase loop v3: - introduce "inline" root node that is embedded into tree object and in which the 1st key is inserted. This greatly improves the 1-key-tree performance, which is pretty common case for rows cache v2: - introduce "linear" root leaf that grows on demand This improves the memory consumption for small trees. This linear node may and should over-grow the NodeSize parameter. This comes from the fact that there are two big per-key memory spikes on small trees -- 1-key root leaf and the first split, when the tree becomes 1-key root with two half-filled leaves. If the linear extention goes above NodeSize it can flatten even the 2nd peak - mitigate the keys indirection a bit Prefetching the keys while doing the intra-node linear scan and the nodes while descending the tree gives ~+5% of fill and find - generalize stress tests for B and B+ trees - cosmetic changes TODO: - fix few inefficincies in the core code (walks the sub-tree twice sometimes) - try to optimize the leaf nodes, that are not lef-/righmost not to carry unused tree pointer on board Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2021-02-02 09:30:29 +03:00
Avi Kivity	32cdcc0c8b	Merge "sstables: consolidate reader factory methods" from Botond " Currently there are three different methods for creating an sstable reader: * one for single key reads * one for ranged reads * and one nobody uses This patch-set consolidates all these into a single `make_reader()` method, which behind the scenes uses the same logic to dispatch to the right sstable reader constructor that `sstables::as_mutation_source()` uses. This patch-set is part of an effort to clean up the jungle that is the various reader creation methods. The next step is to clean up the sstable_set, which has even more methods. One very sad discovery I made while working on this patch-set is that we still default `mutation_reader::forwarding` to `yes` in the sstable range reader creator method and in the `mutation_source::make_reader()`. I couldn't assume that all callers are passing what they mean as the value for that parameter. I found many sites in tests that create forwardable single partition readers. This is also something we should address soon. Tests: unit(release, debug:v3) " * 'sstables-consolidate-reader-factory-methods-v4' of https://github.com/denesb/scylla: cql_query_test: add unit test covering the non-optimal TWCS sstable read path sstable_mutation_reader: consolidate constructors tests: don't pass temporary ranges to readers sstables: sstable_mutation_reader: remove now unused whole sstable constructor sstables: stats: remove now unused sstable_partition_reads counter sstable: remove read_.row._flat() methods tree-wide: use sstables::make_reader() instead of the read_.row._flat() methods sstables: pass partition_range to create_single_key_sstable_reader() sstables: sstable: add make_reader()	2021-01-28 12:05:06 +02:00
Konstantin Osipov	b4f875f08e	uuid: reduce code dependency on UUID_gen.hh Do not include UUID_gen.hh in trace_state.hh and lists.hh to reduce header level dependency on it. Message-Id: <20210127173114.725761-2-kostja@scylladb.com>	2021-01-27 20:08:29 +02:00
Botond Dénes	c3b4e990a2	tree-wide: use sstables::make_reader() instead of the read_.row._flat() methods	2021-01-27 17:38:17 +02:00
Avi Kivity	aec231ba2e	Merge "Unify query paths" from Botond " Currently we have two parallel query paths: * database::query() -> table::query() -> data_query() * mutation::query() The former is used by single partition queries, the latter by range scans, as mutation::query() is used to convert reconcilable_result to query::result (which means it is also used in single partition queries if it triggers read repair). This is a rather unfortunate situation as we have two parallel implementation of the query code, which means they are prone to diverge, and in fact they already have -- more on that later. This patchset aims to remedy this situation by retiring `mutation::query()` and migrating users to an implementation based on the "standard" query path, in other words one using the same building blocks as the `database::query()` path. This means using `compact_mutation` for compacting and `query_result_builder` for result building. These components however were created to work with `flat_mutation_reader`, however introducing a reader into this pipeline would mean that we'd have to make all the related APIs asynchronous, which would cause an insane amount of churn. To avoid this, this patchset adds an API compatible `consume()` method to `mutation`, which can accept a `compact_mutation` instance as-is. This allows an elegant and succinct reimplementation. So far so good. Like mentioned above, the two implementations have diverged in time, or have been different from the start. The difference manifest when calculating digests, more precisely in which tombstones are included in the digest. The retired `mutation::query()` path incorporates only non-purgeable tombstones in the digest. The standard query path however incorporates all tombstones, even those that can be purged. After some scrutiny however this difference proved to be completely theoretical, as the code path where this would matter -- converting reconcilable result to query result -- passes min timestamp as the query time to the compaction, so nothing is compacted and hence the difference has no chance to manifest. This patch-set was motivated by the desire to provide a single solution to #7434, instead of two, one for each path. Tests: unit(release:v2, debug:v2, dev:v3) " * 'unified-query-path/v3' of https://github.com/denesb/scylla: mutation: remove now unused query() and query_compacted() treewide: use query_mutations() instead of mutation::query() mutation_test: test_query_digest: ensure digest is produced consistently mutation_query: introduce query_mutation() mutation_query: to_data_query_result(): migrate to standard query code mutation_query: move to_data_query_result() to mutation_partition.cc mutation: add consume() flat_mutation_reader: move mutation consumer concepts to separate header mutation compactor: query compaction: ignore purgeable tombstones	2021-01-27 15:58:47 +02:00
Benny Halevy	1847d49971	test: test_env: pick the highest sstable version by default If possible, test the highest sstable format version, as it's the mostly used. If there pre-written sstables we need to load from the test directory from an older version, either specify their version explicitly, or use the new test_env::reusable_sst method that looks up the latest sstable version in the given directory and generation. Test: unit(release) Signed-off-by: Benny Halevy <bhalevy@scylladb.com> Message-Id: <20201210161822.2833510-1-bhalevy@scylladb.com>	2021-01-24 10:38:55 +02:00
Botond Dénes	1a3ee71b39	treewide: use query_mutations() instead of mutation::query() We want to retire the latter.	2021-01-22 15:36:37 +02:00
Benny Halevy	29002e3b48	flat_mutation_reader: return future from next_partition To allow it to asynchronously close underlying readers on next_partition(). Signed-off-by: Benny Halevy <bhalevy@scylladb.com>	2021-01-13 17:35:07 +02:00
Raphael S. Carvalho	198b87503f	row_cache: allow external updater to decouple preparation from execution External updater may do some preparatory work like constructing a new sstable list, and at the end atomically replace the old list by the new one. Decoupling the preparation from execution will give us the following benefits: - the preparation step can now yield if needed to avoid reactor stalls, as it's been futurized. - the execution step will now be able to provide strong exception guarantees, as it's now decoupled from the preparation step which can be non-exception-safe. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2020-12-28 13:17:45 -03:00

1 2 3

117 Commits