scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-24 10:30:38 +00:00

Author	SHA1	Message	Date
Avi Kivity	a95d3f9cf5	Merge "Commitlog shutdown" from Calle "Refs #293 * Add a commitlog::sync_all_segments, that explicitly forces all pending disk writes * Only delete segments from disk IFF they are marked clean. Thus on partial shutdown or whatnot, even if CL is destroyed (destructor runs) disk files not yet clean visavi sstables are preserved and replayable * Do a sync_all_segments first of all in database::stop. Exactly what to not stop in main I leave up to others discretion, or at least another patch."	2015-09-08 11:11:18 +03:00
Tomasz Grabiec	15ae1a92cb	Merge branch 'pdziepak/compaction-remove-items/v4' from seastar-dev.git From Pawel: This series makes compaction remove items that are no longer items: - expired cells are changed into tombstones - items covered by higher level tombstones are removed - expired tombstones are removed if possible Fixes #70. Fixes #71.	2015-09-08 09:23:00 +02:00
Paweł Dziepak	64949e8339	schema: make gc_grace_seconds() return gc_clock::duration Signed-off-by: Paweł Dziepak <pdziepak@cloudius-systems.com>	2015-09-07 21:14:41 +02:00
Calle Wilund	256c0550bf	Commitlog: Only delete segments on disk if they are marked clean For #293 - i.e. allow more or less coherent shutdown/destruction of the commitlog while retaining disk data. (tests still clear stuff explicitly).	2015-09-07 20:32:01 +02:00
Calle Wilund	4ed95b7020	Commitlog: Add sync_all_segments() For #293 - allows explicit flush to disk (not close!) of all active segments	2015-09-07 20:31:59 +02:00
Calle Wilund	d614143f5e	Commitlog/database: Fixup series "Commit log flush request on disk overflow" Also at seastar-dev: calle/commitlog_flush_v3 (And, yes, this time I _did_ update the remote!) Refs #262 Commit of original series was done on stale version (v2) due to authors inability to multitask and update git repos. v3: * Removed future<> return value from callbacks. I.e. flush callback is now only fully syncronous over actual call	2015-09-07 21:29:19 +03:00
Avi Kivity	dee9060b12	Merge "Commit log flush request on disk overflow" from Calle "Fixes #262 Handles CL disk size exceeding configured max size by calling flush handlers for each dirty CF id / high replay_position mark. (Instead of uncontrolled delete as previously). * Increased default max disk size to 8GB. Same as Origin/scylla.yaml (so no real change, but synced). * Divide the max disk size by cpus (so sum of all shards == max) * Abstract flush callbacks in CL * Handler in DB that initiates memtable->sstable writes when called. Note that the flush request is done "syncronously" in new_segment() (i.e. when getting a new segment and crossing threshold). This is however more or less congruent with Origin, which will do a request-sync in the corresponding case. Actual dealing with the request should at least in production code however be done async, and in DB it is, i.e. we initiate sstable writes. Hopefully they finish soon, and CL segments will be released (before next segment is allocated). If the flush request does _not_ eventually result in any CF:s becoming clean and segments released we could potentially be issuing flushes repeatedly, but never more often than on every new segment."	2015-09-07 18:46:48 +03:00
Gleb Natapov	da242146b6	do not pass storage_proxy reference across cpus storage_proxy instances are per cpu, so they cannot be passed around to other cpus.	2015-09-07 17:16:29 +02:00
Calle Wilund	fdb921afb2	Commitlog: Add flushing of segment CF:s on disk overflow * Do not throw away commitlog segments on disk size overflow. Issue a flush request (i.e. calculate RP we want to free unto, and for all dirty CF:s, do a request). "Abstracted" as registerable callback. I.e. DB:s responsibility to actually do something with it.	2015-09-07 13:21:43 +02:00
Calle Wilund	31f2dcb342	Config: change commilog max size on disk to be in sync with scylla.yaml	2015-09-07 13:13:51 +02:00
Calle Wilund	841dd32a8a	Commitlog: divide max on-disk-size by num cpus To try to keep the resulting limit as configured	2015-09-07 13:13:46 +02:00
Asias He	f89a25562c	storage_service: Fix is_auto_bootstrap Get the value from cfg option.	2015-09-07 12:53:58 +03:00
Asias He	7cc768a864	gossip: Fix wrong cluster name and partitioner name Right now, gossip returns hard coded cluster and partitioner name. sstring get_cluster_name() { // FIXME: DatabaseDescriptor.getClusterName() return "my_cluster_name"; } sstring get_partitioner_name() { // FIXME: DatabaseDescriptor.getPartitionerName() return "my_partitioner_name"; } Fix it by setting the correct name from configure option. With this cqlsh 127.0.0.$i -e "SELECT * from system.local; returns correct cluster_name. Fixes #291	2015-09-07 09:21:18 +03:00
Avi Kivity	42dc29619d	Merge "Optimize mutation copies and moves" from Paweł "This series deals with copies and moves of mutation. The former are dealt with by adding std::move() and missing 'mutable' (in case of lambdas). The latter are improved by storing mutation_partition externally thus removing the need for moving mutation_partition each time mutation is moved. Storing mutation_partition externally is obviously trading the cost of move constructor for the cost of allocation which shows in perf_mutation results since mutations aren't moved in that test. perf_mutation (-c 1): before: 3289520.06 tps after: 3183023.37 tps diff: -3.24% perf_simple_query (read): before: 526954.05 tps after: 577225.16 tps diff +9.54% perf_simple_query (write): before: 731832.70 tps after: 734923.60 tps diff: +0.42% Fixes #150 (well, not completely)."	2015-09-03 12:05:28 +03:00
Paweł Dziepak	830a86258b	db: avoid copying mutations Signed-off-by: Paweł Dziepak <pdziepak@cloudius-systems.com>	2015-09-03 10:30:32 +02:00
Paweł Dziepak	ddec2b4d09	batchlog_manager: pass mutations by const ref Signed-off-by: Paweł Dziepak <pdziepak@cloudius-systems.com>	2015-09-03 10:30:29 +02:00
Paweł Dziepak	8188896eb7	schema_tables: add missing mutable Signed-off-by: Paweł Dziepak <pdziepak@cloudius-systems.com>	2015-09-03 10:30:25 +02:00
Glauber Costa	28f315fad4	system_keyspace: keep msg alive when needed Fixes #266 Some callsites are fine: if we just get the message and process it, as is the case with check_health for instance, msg will be alive and all is good. But if we return a future inside the processing, msg must be kept alive. Classic bug, appearing again. Pekka saw this in practice in another bug. We haven't seen anything that is related to this, but it is certainly wrong. Signed-off-by: Glauber Costa <glommer@cloudius-systems.com>	2015-09-03 09:11:07 +03:00
Pekka Enberg	ce39f9d57a	db/system_keyspace: Fix use-after-free in build_dc_rack_info() Fixes #264. Signed-off-by: Pekka Enberg <penberg@cloudius-systems.com>	2015-09-02 16:37:34 +03:00
Calle Wilund	d95101664d	Commitlog: Don't throw exceptions on unrecognized files in CL dir	2015-09-01 14:23:03 +02:00
Calle Wilund	1814f89730	Commitlog: Add some more metrics + accessors for json API Fixes #99 Adding missing commitlog metrics to the rest API. v2: Mis-send (clumsy fingers) v3: Use map_reduce0 + subroutine for nicer code v4: rebased on current master v5: rebased yet again. Since the _second_ file in this previous patch set was commited, and is dependent on this very change below to even compile, some expediency might be warranted.	2015-09-01 10:15:33 +03:00
Calle Wilund	9ba84e458a	Commitlog: Handle partial writes in segment::cycle * Fixes #247 * Re-introduce test_allocation_failure, but allow for the "failure" to not happen. I.e. if run with low memory settings, the test will check that allocation failure is graceful. With lots of memory it will check partial write.	2015-08-31 20:02:05 +03:00
Calle Wilund	d3a01072af	CommitLogReplayer: Java -> C++ Initial implementation	2015-08-31 14:29:50 +02:00
Calle Wilund	bbf82e80d0	Commitlog: Allow skipping X bytes in commit log reader Also refactor reader into named methods for debugging sanity.	2015-08-31 14:29:49 +02:00
Calle Wilund	da9ea641e5	Commitlog: Handle full paths in descriptor file name parse.	2015-08-31 14:29:48 +02:00
Calle Wilund	02d2bef1f2	Commitlog: Expose convinience method "list_existing_segments"	2015-08-31 14:29:48 +02:00
Calle Wilund	19052b3c09	Commitlog: Expose list_existing_descriptors	2015-08-31 14:29:48 +02:00
Calle Wilund	e068ffb5a5	Commitlog: Make file reader provide replay_position for entries	2015-08-31 14:29:47 +02:00
Calle Wilund	41b1ad8600	Commitlog: Make descriptor type visible/usable from outside	2015-08-31 14:29:47 +02:00
Calle Wilund	ea38b223bd	Commitlog: change the ID generation scheme * Make it more like origin, i.e. based on wall clock time of app start * Encode shard ID in the, RP segement ID, to ensure RP:s and segement names are unique per shard	2015-08-31 14:29:46 +02:00
Calle Wilund	0fcf7e3e91	Commitlog: Make "position" type 32-bit to align replay_position with Origin * Note: removed commitlog_test:test_allocation_failure because with segments limited to 4GB -> mutation limited to 2GB, actually forcing a fail is not guaranteed or even likely.	2015-08-31 14:29:44 +02:00
Calle Wilund	3f1a91b89c	Commitlog: do not eagerly create first segment on init Deferring makes it easier to separate old segments from new, which in turn helps replay logic.	2015-08-31 13:11:44 +02:00
Avi Kivity	5f62f7a288	Revert "Merge "Commit log replay" from Calle" Due to test breakage. This reverts commit `43a4491043`, reversing changes made to `5dcf1ab71a`.	2015-08-27 12:39:08 +03:00
Asias He	80c996a315	db/system_keyspace: Fix get_local_host_id Before: host_id in system.local is empty After: host_id in system.local is inserted correctly This fixes a hasty problem that we always get a new host_id when booting up a node with data.	2015-08-27 11:01:07 +03:00
Avi Kivity	43a4491043	Merge "Commit log replay" from Calle "Initial implementation/transposition of commit log replay. * Changes replay position to be shard aware * Commit log segment ID:s now follow basically the same scheme as origin; max(previous ID, wall clock time in ms) + shard info (for us) * SStables now use the DB definition of replay_position. * Stores and propagates (compaction) flush replay positions in sstables * If CL segments are left over from a previous run, they, and existing sstables are inspected for high water mark, and then replayed from those marks to amend mutations potentially lost in a crash * Note that CPU count change is "handled" in so much that shard matching is per _previous_ runs shards, not current. Known limitations: * Mutations deserialized from old CL segments are _not_ fully validated against existing schemas. * System::truncated_at (not currently used) does not handle sharding afaik, so watermark ID:s coming from there are dubious. * Mutations that fail to apply (invalid, broken) are not placed in blob files like origin. Partly because I am lazy, but also partly because our serial format differs, and we currently have no tools to do anything useful with it * No replay filtering (Origin allows a system property to designate a filter file, detailing which keyspace/cf:s to replay). Partly because we have no system properties. There is no unit test for the commit log replayer (yet). Because I could not really come up with a good one given the test infrastructure that exists (tricky to kill stuff just "right"). The functionality is verified by manual testing, i.e. running scylla, building up data (cassandra-stress), kill -9 + restart. This of course does not really fully validate whether the resulting DB is 100% valid compared to the one at k-9, but at least it verified that replay took place, and mutations where applied. (Note that origin also lacks validity testing)"	2015-08-27 10:53:36 +03:00
Glauber Costa	391eea564e	system_tables: implement load_host_id A simple translation from the original code. Signed-off-by: Glauber Costa <glommer@cloudius-systems.com>	2015-08-25 19:16:30 -05:00
Glauber Costa	0fd2861293	system_tables: implement load_tokens A simple translation from the original code Signed-off-by: Glauber Costa <glommer@cloudius-systems.com>	2015-08-25 19:16:30 -05:00
Calle Wilund	2a1c7d2587	CommitLogReplayer: Java -> C++ Initial implementation	2015-08-25 09:41:56 +02:00
Calle Wilund	86a97fea4c	Commitlog: Allow skipping X bytes in commit log reader Also refactor reader into named methods for debugging sanity.	2015-08-25 09:41:55 +02:00
Calle Wilund	37cfc09e91	Commitlog: Handle full paths in descriptor file name parse.	2015-08-25 09:41:55 +02:00
Calle Wilund	4364d72ca3	Commitlog: Expose convinience method "list_existing_segments"	2015-08-25 09:41:54 +02:00
Calle Wilund	a3a02968ab	Commitlog: Expose list_existing_descriptors	2015-08-25 09:41:54 +02:00
Calle Wilund	fcb87471b9	Commitlog: Make file reader provide replay_position for entries	2015-08-25 09:40:53 +02:00
Calle Wilund	db6370ad87	Commitlog: Make descriptor type visible/usable from outside	2015-08-25 09:40:53 +02:00
Calle Wilund	4f24b9795e	Commitlog: change the ID generation scheme * Make it more like origin, i.e. based on wall clock time of app start * Encode shard ID in the, RP segement ID, to ensure RP:s and segement names are unique per shard	2015-08-25 09:40:52 +02:00
Asias He	22ee468428	db/system_keyspace: Fix set_bootstrap_state We set status to COMPLETED in join_token_ring set_bootstrap_state(db::system_keyspace::bootstrap_state::COMPLETED) but cqlsh 127.0.0.$i -e "SELECT * from system.local;" shows bootstrapped -> IN_PROGRESS The static sstring state_name is the bad boy.	2015-08-24 18:54:42 +08:00
Avi Kivity	8a4648761c	tests: make test cql environment use volatile system keyspace Prevents hangs due to the database not being able to persist a memtable. Tested-by: Asias He <asias@cloudius-systems.com>	2015-08-24 13:50:22 +03:00
Pekka Enberg	9476e3f19f	db/index: Kill Java code Signed-off-by: Pekka Enberg <penberg@cloudius-systems.com>	2015-08-24 11:51:50 +03:00
Pekka Enberg	544c7936d8	db/commitlog: Kill Java code Signed-off-by: Pekka Enberg <penberg@cloudius-systems.com>	2015-08-24 11:51:49 +03:00
Calle Wilund	ac74dd6159	Commitlog: Make "position" type 32-bit to align replay_position with Origin	2015-08-24 10:05:44 +02:00

1 2 3 4 5 ...

403 Commits