scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-24 18:40:38 +00:00

Author	SHA1	Message	Date
Gleb Natapov	f5679e0416	database: remove remnants of no longer existing db::serializer. Message-Id: <20170604100552.GD8248@scylladb.com>	2017-06-04 13:07:17 +03:00
Raphael S. Carvalho	4b4a1883aa	refresh: do not use default priority for loading new sstables Metadata is read using default priority class, which can significantly slow down the process under high load. Compaction class can be used, and if it turns out to be a problem, we can switch to a special class for it. Fixes #1859. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Message-Id: <20170517184546.17497-1-raphaelsc@scylladb.com>	2017-05-22 19:03:17 +03:00
Raphael S. Carvalho	a4e414cb3b	database: reduce memory requirement to load sstables SSTable load temporarily uses more space than needed to store metadata, due to: 1) All components are read using read_simple() which uses 128k buffer. file::dma_read_bulk() will allocate 128k, and may potentially allocate another big buffer (128k - read) for file::read_maybe_eof(). 2) read_filter() may use double the space it needs to. Due to the fact that sstable loading parallelism is unlimited, Scylla may require much more memory to load all sstables, and that may lead to OOM. Higher the number of sstables higher the memory overhead. To confirm this problem, I wrote a test[1] which loads 30k sstables in parallel and reports the memory usage peak in the end. When loading 30k sstables, each of which metadata is ~300kb, memory usage peak was ~18G. When loading completed, only ~9GB were needed to store all the metadata. [1]: https://gist.github.com/raphaelsc/2db37b4fb34301833ab9eeed3b1a524d To fix this problem, we need to set a limit on load parallelism (let's start with a small number like 3 and adjust later if needed) and rely on readahead so that the requirement drops considerably without increasing boot time. Actually, boot time is improved by it. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Reviewed-by: Nadav Har'El <nyh@scylladb.com>	2017-05-22 11:52:51 -03:00
Duarte Nunes	983af595e9	database: Read existing base mutations When generating updates for a materialized view we need to read the existing base row, to be able to determine the primary key of the view row the new base update will supplant, in case the view includes a base non-primary key column in its own primary key. That old view row will be tombstoned or updated, if it exists, depending on the difference between the new base row and the existing one, if any. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-05-17 10:33:19 +02:00
Calle Wilund	3514123677	database: Extract "remove" from "drop_columnfamily"	2017-05-10 16:44:48 +00:00
Calle Wilund	48ddcbb77b	database/main: encapsulate system CF dir touching	2017-05-10 16:44:47 +00:00
Avi Kivity	8c5c5d3004	Merge "CQL front-end for secondary indices" from Pekka "This patch series adds CQL front-end support for secondary indices. You can now execute CREATE INDEX and DROP INDEX statements, which will update the newly added "Indexes" system table. However, the indexes are not actually backed up by anything nor are they available for CQL queries. The feature is hidden behind a new cluster feature flag and enabled only with the "--experimental" flag." * 'penberg/cql-2i/v2' of github.com:cloudius-systems/seastar-dev: (34 commits) schema: Kill index_type enum schema: Kill index_info class cql3/statements/create_index_statement: Use database::existing_index_names() in validation cql3/statements: Use secondary index manager in alter_table_statement class index: Add secondary_index_manager thrift/handler: Use index_metadata db/schema_tables: Index persistence schema: Add all_indices() to schema class schema: Remove add_default_index_names() from schema_builder class db/schema_tables: Add system table for indices cql3/Cgl.g: DROP INDEX cql3/statements: Add drop_index_statement class database: Add find_indexed_table() to database class cql3: Return change event from announce_migration() cql3/statements: Multiple index targets for CREATE INDEX cql3/statements: Use index_metadata in create_index_statement class cql3/statements: Use feature flag in create_index_statement class service/storage_service: Add feature flag for secondary indices database: Add get_available_index_name() to database class schema: Add get_default_index_name() to index_metadata class ...	2017-05-08 17:04:40 +03:00
Pekka Enberg	f26b8d7afb	database: Add find_indexed_table() to database class	2017-05-04 14:59:12 +03:00
Pekka Enberg	930fa79aff	database: Add get_available_index_name() to database class	2017-05-04 14:59:11 +03:00
Pekka Enberg	c6e7d4484a	database: Make existing_index_names() per-keyspace operation	2017-05-04 14:59:11 +03:00
Paweł Dziepak	24f4dcf9e4	db: make virtual dirty soft limit configurable Message-Id: <20170428150005.28454-1-pdziepak@scylladb.com>	2017-04-30 19:17:22 +03:00
Raphael S. Carvalho	662fe77c11	database: kill column_family::start_rewrite Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:33 -03:00
Raphael S. Carvalho	cf45333588	database: implement new sstable resharding algorithm NOTE: it's not wired yet. Currently, a shared sstable is rewritten at all shards it belongs to and only after that, it's deleted. With this new algorithm, a shared sstable will be read only once and N unshared sstables will be created, each of them with 1/N of the data. After it's done, each owner shard will receive its new unshared sstable replacing its ancestors. Another benefit is that we'll no longer have resharding resulting in number of sstables growing considerably after resharding. A full-sized leveled sstable is usually 160MB, so after resharding, we could have N files of 160MB/N. Now, leveled strategy will help resharding. N adjacent sstables of same level will be resharded together, so we'll end up with N files of N*160MB/N. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:30 -03:00
Raphael S. Carvalho	6513252e91	database: introduce function to replace new sstables by their ancestors When resharding, we're working with sstables from all shards. So let's say we're done with resharding of sstable A that belongs to shard 0 and 1 and sstable B that belongs to shard 1 and 2. SStables were generated for shards 0, 1, and 2. So shards 0, 1, and 2 need to load the new sstables and remove the ancestors. Shard 1 for example will remove sstables A and B (ancestors) and add the new one. Then it comes this new function. We'll forward new sstables to their target shards using foreign sstable open info. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:27 -03:00
Raphael S. Carvalho	c44a2319e6	prevent regular compaction from choosing shared sstables For new resharding, it's important to exclude resharding sstables from the list of candidates for regular compaction. That's doesn't affect current resharding because it marks the sstables as compacting. That won't work with new resharding which will work with sstables from multiple shards. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:26 -03:00
Raphael S. Carvalho	405e41e9a8	database: export column family dir Reviewed-by: Nadav Har'El <nyh@scylladb.com> Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:19 -03:00
Raphael S. Carvalho	2b774c5bc3	database: inform if column family has shared tables That's gonna be useful to quickly determine if it's worth resharding a column family. Reviewed-by: Nadav Har'El <nyh@scylladb.com> Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com>	2017-04-21 17:11:17 -03:00
Amnon Heiman	7572addfbf	column_family: metrics should be register once column_family constructor uses delegation, as such, only the actual constructor implementation should contain a call to register the metrics. Current implementation ends up with re registration of the metrics. Signed-off-by: Amnon Heiman <amnon@scylladb.com> Message-Id: <20170320140817.22214-1-amnon@scylladb.com>	2017-03-21 11:20:29 +02:00
Duarte Nunes	876a514743	database: Upgrade mutation to current schema to push view updates This patch ensures we upgrade the mutation to the current schema when generating and pushing view updates, so that the it matches the most up to date views. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-03-15 18:15:27 +01:00
Duarte Nunes	bfb8a3c172	materialized views: Replace db::view::view class The write path uses a base schema at a particular version, and we want it to use the materialized views at the corresponding version. To achieve this, we need to map the state currently in db::view::view to a particular schema version, which this patch does by introducing the view_info class to hold the state previously in db::view::view, and by having a view schema directly point to it. The changes in the patch are thus: 1) Introduce view_info to hold the extra view state; 2) Point to the view_info from the schema; 3) Make the functions in the now stateless db::view::view non-member; 4) Remove the db::view::view class. All changes are structural and don't affect current behavior. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-03-15 15:50:05 +01:00
Paweł Dziepak	38c1501f4d	db: make apply an execution stage	2017-03-09 09:27:43 +00:00
Avi Kivity	439b38f5ab	Merge "Improvements to counter implementation" from Paweł "This series adds various optimisations to counter implementation (nothing extreme, mostly just avoiding unnecessary operations) as well as some missing features such as tracing and dropping timed out queries. Performance was tested using: perf-simple-query -c4 --counters --duration 60 The following results are medians. before after diff write 18640.41 33156.81 +77.9% read 58002.32 62733.93 +8.2%" * tag 'pdziepak/optimise-counters/v3' of github.com:cloudius-systems/seastar-dev: (30 commits) cell_locker: add metrics for lock acquisition storage_proxy: count counter updates for which the node was a leader storage_proxy: use counter-specific timeout for writes storage_proxy: transform counter timeouts to mutation_write_timeout_exception db: avoid allocations in do_apply_counter_update() tests/counters: add test for apply reversability counters: attempt to apply in place atomic_cell: add COUNTER_IN_PLACE_REVERT flag counters: add equality operators counters: implement decrement operators for shard_iterator counters: allow using both views and mutable_views atomic_cell: introduce atomic_cell_mutable_view managed_bytes: add cast to mutable_view bytes: add bytes_mutable_view utils: introduce mutable_view db: add more tracing events for counter writes db: propagate tracing state for counter writes tests/cell_locker: add test for timing out lock acquisition counter_cell_locker: allow setting timeouts db: propagate timeout for counter writes ...	2017-03-07 11:48:13 +02:00
Avi Kivity	1af9e3a5cb	Merge "database: fix the 'nodetool clearsnapshot'" from Vlad "Work on this series started with fixing the 'nodetool clearsnapshot'. The current master code ignores the snapshots in deleted keyspaces (issue #2045). I noticed that in many places our code has to build the path to some directory/file it simply had the sstring(<path1>) + "/" + sstring(<path2>) constructs which may cause us issues if somebody decides to complile/run scylla on not-Unix-based OS, like Microsoft Windows. I understand that this is a long shot but if we can make it right now - why not to. The answer is boost::filesystem::path class - its synchronous parts, of course. I decided to take an initiative and fix the issues above and then use the fixed code for fixing the issue #2045: - Fix some minor issues in the existing code. - Extend the lister class and move it into the separate files outside database.cc. On the way I've found an issue in the existing code (issue #2071). This series fixes this one too (PATCH2)."	2017-03-06 16:45:31 +02:00
Paweł Dziepak	04b80272f2	cell_locker: add metrics for lock acquisition	2017-03-02 09:05:12 +00:00
Paweł Dziepak	277501f42f	db: propagate tracing state for counter writes	2017-03-02 09:05:10 +00:00
Paweł Dziepak	25173f8095	db: propagate timeout for counter writes	2017-03-02 09:05:10 +00:00
Paweł Dziepak	f25fa6566f	db: avoid deserialization when applying counter mutation In the later stages of counter write path a mutation is produced that already has all cells transformed to counter shards and can be applied to the memtable and written to the commitlog. The current interface expectes a frozen mutation, which is suboptimal for counters. The freeze itself is unaviodable -- it is required by commitlog, but we can avoid later deserialization of frozen_mutation when it is applied to the memtable if we pass the unfrozen mutation along.	2017-03-01 16:33:37 +00:00
Paweł Dziepak	426345e1d4	storage_proxy: avoid excessive mutation freezes	2017-03-01 16:33:36 +00:00
Paweł Dziepak	0198d8e470	Merge "Introduce streamed_mutation::fast_forward_to()" from Tomasz "This introduces an API which allows forward navigation in a stream of mutation fragments. It allows one to consume only a subset of the stream by iteratively specifying sub-ranges from which fragments should be returned. API outline: When in forwarding mode, the stream does not return all fragments right away, but only those belonging to the current range. Initially current range only covers the static row. The stream can be forwarded, even before reaching end- of-stream for current range, to a later range with fast_forward_to(). Forwarding doesn't change initial restrictions of the stream, it can only be used to skip over data. Monotonicity of positions is preserved by forwarding. That is fragments emitted after forwarding will have greater positions than any fragments emitted before forwarding. For any range, all range tombstones relevant for that range which are present in the original stream will be emitted. Range tombstones emitted before forwarding which overlap with the new range are not necessarily re-emitted. When not in forwarding mode, the stream acts as if the current range was equal to the full range. This implies that fast_forward_to() cannot be used. Whether stream is in forwarding mode or not is specified when the stream is created, typically via mutation_source interface. What's left for later series: Optimization by providing specialized implementations. This series implements forwarding support in all mutation sources via generic wrapper which simply drops fragments." * tag 'tgrabiec/clustering-fast-forward-to-v2' of github.com:scylladb/seastar-dev: tests: mutation_source_tests: Verify monotonicty of positions tests: random_mutation_generator: Spread the keys more tests: mutation_source_test: Make blobs more easily distinguishable tests: streamed_mutation: Test that merged stream passes mutation source tests tests: mutation_source_test: Add tests for forwarding of streamed_mutation tests: streamed_mutation_assertions: Add methods for navigating the stream tests: Add range generators to random_mutation_generator partition_slice_builder: Add with_ranges() query: Introduce full_clustering_range streamed_mutation: Add non-owning variant of mutation_from_streamed_mutation() db: Enable creating forwardable readers via mutation_source mutation_source: Document liveness requirements mutation_source: Cleanup db: Replace virtual_reader_type with mutation_source_opt partition_version: Refactor make_partition_snapshot_reader() overloads database: Fix mutation_source created by as_mutation_source() to not ignore trace_state_ptr memtable: Accept all mutation_source parameters streamed_mutation: Implement fast_forward_to() in stream merger streamed_mutation: Add generic implementation of forwardable streamed_mutation streamed_mutation: Add fast_forward_to() API position_in_partition: Introduce position_range position_in_partition: Introduce position constructor for right after the static row streamed_mutation: Make cast to view non-explicit streamed_mutation: Make schema() getter non-copying	2017-02-24 10:37:51 +00:00
Tomasz Grabiec	892d4a2165	db: Enable creating forwardable readers via mutation_source Right now all mutation source implementations will use make_forwardable() wrapper.	2017-02-23 18:50:44 +01:00
Tomasz Grabiec	586dbaa8d3	db: Replace virtual_reader_type with mutation_source_opt Virtual reader is a mutation_source.	2017-02-23 18:23:52 +01:00
Tomasz Grabiec	f46ae8128d	database: Fix mutation_source created by as_mutation_source() to not ignore trace_state_ptr It was using the state passed via as_mutation_source() instead. Let's respect mutation_source contract instead, and use the state passed via mutation_source invocation. Technically just a cleanup. Alse prerequisite for more cleanup.	2017-02-23 18:23:52 +01:00
Paweł Dziepak	359c617821	db: restore call to check_valid_rp() `5a0955e89d` "db: add operations for applying counter updates" merged two column_family::apply() overloads into do_apply() in order to reduce code duplication. Unfortunately, a call to check_valid_rp() didn't survive that change. Message-Id: <20170221133800.30411-1-pdziepak@scylladb.com>	2017-02-21 15:26:04 +01:00
Vlad Zolotarov	978241d473	database: move lister class into separate files Move lister class away from database.cc. This is a preparation for moving it to the seastar library. Signed-off-by: Vlad Zolotarov <vladz@scylladb.com>	2017-02-17 17:50:40 -05:00
Vlad Zolotarov	34cafa71c3	database: make 'clearsnapshot' to delete the snapshots of deleted keyspaces if requested The current implementation of 'nodetool clearsnapshot' command only deletes the snapshots of the keyspaces that are alive at the time the command is issued (issue #2045). This, besides not implementing the spec, prevents users from being able to clear the disk space occupied by snapshots of deleted keyspaces that are no longer needed (e.g. snapshots created when KS is deleted). This patch fixes this issue by making the database::clear_snapshot() scan the data directories looking for the snapshots to be deleted instead of relying on in-memory data structures. This patch makes column_family::clear_snapshot() method not needed any more. Fixes #2045 Signed-off-by: Vlad Zolotarov <vladz@scylladb.com>	2017-02-17 17:50:40 -05:00
Vlad Zolotarov	53532ba5ff	database: lister: pass the parent path object to callbacks Pass a parent directory boost::filesystem::path object to the walker and filter callbacks. Signed-off-by: Vlad Zolotarov <vladz@scylladb.com>	2017-02-17 17:50:37 -05:00
Vlad Zolotarov	b4c970dfc6	database: lister: make the "filter" callback receive directory_entry instead of sstring Filter should get all information that the caller has in hand that may be used for filtering. directory_entry has the following information: - Type of the entry - Its name For the code that used lister filters so far this would be enough, however it's not hard to imagine a filter that may need the parent directory as well. We will add the parent directory path in the follow up patches to make the interface complete. Signed-off-by: Vlad Zolotarov <vladz@scylladb.com>	2017-02-17 17:46:59 -05:00
Avi Kivity	9530bac2d6	Merge "Adding metrics using histogram and labels" from Amnon "This series uses the newly added histogram and label support to add metrics to the storage_proxy and to the column_family. This would add latency and histogram and the missing metrics from column family." * 'amnon/histogram_metrics' of github.com:cloudius-systems/seastar-dev: database: add metrics registration for the coloumn family storage_proxy: add read and write latency histogram estimated_histogram: returns a metrics histogram	2017-02-09 11:40:57 +02:00
Amnon Heiman	292c08f598	database: add metrics registration for the coloumn family This patch adds a metrics registration to the column_family. Using label each column metrics is label with its keyspace and column family name. Signed-off-by: Amnon Heiman <amnon@scylladb.com>	2017-02-06 18:27:01 +02:00
Duarte Nunes	0eca6301d3	database: Apply mutation to views This patch changes the database apply path so that it also generates the mutations for the column family's views and sends them to the paired view replicas. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-02-06 13:37:33 +01:00
Duarte Nunes	4777172348	column_family: Push view replica update This patch adds a function to push updates to the view replicas of a particular base table.	2017-02-06 13:36:45 +01:00
Duarte Nunes	16206e9f15	column_family: Generate view updates This patch adds the generate_view_updates() function to the column_family class, which will use the view_update_builder to generate updates to the column_family's materialized views. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-02-06 13:36:45 +01:00
Duarte Nunes	90cb35db04	column_family: Adds affected_views() function This patch the affected_views() to determine the column family's views a given update affects. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-02-06 13:36:45 +01:00
Duarte Nunes	c35d14e285	column_family: Store a pointer to view Instead of storing the view in the column_family's map of materialized views, store a lw_shared_ptr so that the view can be removed while it is being updated. Signed-off-by: Duarte Nunes <duarte@scylladb.com>	2017-02-06 13:35:30 +01:00
Avi Kivity	7a00dd6985	Merge "Avoid avalanche of tasks after memtable flush" from Tomasz "Before, the logic for releasing writes blocked on dirty worked like this: 1) When region group size changes and it is not under pressure and there are some requests blocked, then schedule request releasing task 2) request releasing task, if no pressure, runs one request and if there are still blocked requests, schedules next request releasing task If requests don't change the size of the region group, then either some request executes or there is a request releasing task scheduled. The amount of scheduled tasks is at most 1, there is a single releasing thread. However, if requests themselves would change the size of the group, then each such change would schedule yet another request releasing thread, growing the task queue size by one. The group size can also change when memory is reclaimed from the groups (e.g. when contains sparse segments). Compaction may start many request releasing threads due to group size updates. Such behavior is detrimental for performance and stability if there are a lot of blocked requests. This can happen on 1.5 even with modest concurrency because timed out requests stay in the queue. This is less likely on 1.6 where they are dropped from the queue. The releasing of tasks may start to dominate over other processes in the system. When the amount of scheduled tasks reaches 1000, polling stops and server becomes unresponsive until all of the released requests are done, which is either when they start to block on dirty memory again or run out of blocked requests. It may take a while to reach pressure condition after memtable flush if it brings virtual dirty much below the threshold, which is currently the case for workloads with overwrites producing sparse regions. I saw this happening in a write workload from issue #2021 where the number of request releasing threads grew into thousands. Fix by ensuring there is at most one request releasing thread at a time. There will be one releasing fiber per region group which is woken up when pressure is lifted. It executes blocked requests until pressure occurs." * tag 'tgrabiec/lsa-single-threaded-releasing-v2' of github.com:cloudius-systems/seastar-dev: tests: lsa: Add test for reclaimer starting and stopping tests: lsa: Add request releasing stress test lsa: Avoid avalanche releasing of requests lsa: Move definitions to .cc lsa: Simplify hard pressure notification management lsa: Do not start or stop reclaiming on hard pressure tests: lsa: Adjust to take into account that reclaimers are run synchronously lsa: Document and annotate reclaimer notification callbacks tests: lsa: Use with_timeout() in quiesce()	2017-02-02 17:49:31 +02:00
Paweł Dziepak	5a0955e89d	db: add operations for applying counter updates	2017-02-02 10:35:14 +00:00
Amnon Heiman	45b6070832	Merge seastar upstream * seastar 397685c...c1dbd89 (13): > lowres_clock: drop cache-line alignment for _timer > net/packet: add missing include > Merge "Adding histogram and description support" from Amnon > reactor: Fix the error: cannot bind 'std::unique_ptr' lvalue to 'std::unique_ptr&&' > Set the option '--server' of tests/tcp_sctp_client to be required > core/memory: Remove superfluous assignment > core/memory: Remove dead code > core/reactor: Use logger instead of cerr > fix inverted logic in overprovision parameter > rpc: fix timeout checking condition > rpc: use lowres_clock instead of high resolution one > semaphore: make semaphore's clock configurable > rpc: detect timedout outgoing packets earlier Includes treewide change to accomodate rpc changing its timeout clock to lowres_clock. Includes fixup from Amnon: collectd api should use the metrics getters As part of a preperation of the change in the metrics layer, this change the way the collectd api uses the metrics value to use the getters instead of calling the member directly. This will be important when the internal implementation will changed from union to variant. Signed-off-by: Amnon Heiman <amnon@scylladb.com> Message-Id: <1485457657-17634-1-git-send-email-amnon@scylladb.com>	2017-02-01 14:39:08 +02:00
Tomasz Grabiec	ed9ff19467	lsa: Document and annotate reclaimer notification callbacks They are called from region_group::update(), so must be alloc-free and noexcept.	2017-01-30 19:18:07 +01:00
Raphael S. Carvalho	1857ba0abc	db: fix bad resource usage distribution when resharding due to refresh That's because a single shard is used to calculate generation for new sstables in upload directory, and that will result in that single shard sharing all the resources with other shards. For refresh without upload dir, it currently works fine because we reshuffle column family dir instead. flush_upload_dir() is now a free function, takes a distributed database object, and uses calculate_shard_from_sstable_generation() to decide which shard will move sstable using its own generation namespace. Fixes #2008. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> Message-Id: <b0cccf7bbb61416ff8718bac92fdca90cc5fb9c9.1484253232.git.raphaelsc@scylladb.com>	2017-01-19 18:55:21 +02:00
Duarte Nunes	d53f96e0da	column_family: Only update stats once for a shared sstables This patch ensures that when adding a shared sstable, we select only one cpu to update that column family's stats. This is important so we don't overestimated the on-disk size of sstables when resharding This fixes only a temporary miscount of the current load, since shared sstables are eventually re-written, but a fixes a permanent miscount of the total load. Refs #1592 Signed-off-by: Duarte Nunes <duarte@scylladb.com> Message-Id: <20170119144823.31041-1-duarte@scylladb.com>	2017-01-19 17:40:35 +02:00

1 2 3 4 5 ...

509 Commits