Commit Graph
2307 Commits
Author SHA1 Message Date
Zach Brown c678923401 scoutfs: don't try to sync on mount errors
kill_sb tries to sync before calling kill_block_super.   It shouldn't do
this on mount errors that wouldn't have initialized the higher level
systems needed for syncing.

Signed-off-by: Zach Brown <zab@versity.com>
2017-05-16 10:48:12 -07:00
Zach Brown 66dd35b9a5 scoutfs: fix ring next/prev
The ring node rb walker was returning an exact match for the search key
instead of the last node that was traversed.  This stopped callers from
then iterating from the traversed node to find the next or previous
node.

Signed-off-by: Zach Brown <zab@versity.com>
2017-05-16 10:48:11 -07:00
Zach Brown 723e0368f8 scoutfs: add a trace point for item insertion
Signed-off-by: Zach Brown <zab@versity.com>
2017-05-16 10:48:11 -07:00
Zach Brown 6afeb97802 scoutfs: reference file data with extent items
Our first attempt at storing file data put them in items.  This was easy
to implement but won't be acceptable in the long term.  The cost of the
power of LSM indexing is compaction overhead.  That's acceptable for
fine grained metadata but is totally unacceptable for bulk file data.

This switches to storing file data in seperate block allocations which
are referenced by extent items.

The bulk of the change is the mechanics of working with extents.  We
have high level callers which add or remove logical extents and then
underlying mechanisms that insert, merge, or split the items that
the extents are stored in.

We have three types of extent items.  The primary type maps logical file
regions to physical block extents.  The next two store free extents
per-node so that clients don't create lock and LSM contention as they
try and allocate extents.

To fill those per-node free extents we add messages that communcate free
extents in the form of lists of segment allocations from the server.

We don't do any fancy multi-block allocation yet.  We only allocate
blocks in get_blocks as writes find unmapped blocks.  We do use some
per-task cursors to cache block allocation positions so that these
single block allocations are very likely to merge into larger extents as
tasks stream wites.

This is just the first chunk of the extent work that's coming.  A later
patch adds offline flags and fixes up the change nonsense that seemed
like a good idea here.

The final moving part is that we initiate writeback on all newly
allocated extents before we commit the metadata that references the new
blocks.  We do this with our own dirty inode tracking because the high
level vfs methods are unusably slow in some upstream kernels (they walk
all inodes, not just dirty inodes.)

Signed-off-by: Zach Brown <zab@versity.com>
2017-05-16 10:48:11 -07:00
Zach Brown 6719733ddc scoutfs: output full dirent name when tracing
The dirent name formatting code accidentally copied the calculation for
the length of the name from the xattrs, which are null terminated.  The
durents are not, their length is just the value length minus the dirent
header.

Signed-off-by: Zach Brown <zab@versity.com>
2017-05-16 10:29:59 -07:00
Zach Brown d5a2b0a6db Move towards compaction messages
The compaction code is still directly referencing the super block
and calling sync methods as though it was still standalone.  This is
mostly OK because only the server runs it.  But it isn't quite right
because the sync methods no longer make the rings persistent as they
write the item transaction.  The server is in control of that now.

Eventually we'll have compaction messages being sent between the mount
clients and the server.  Let's take a step in that direction by having
the compaction work call net methods to get its compaction parameters
and finish the compaction.  Eventually these would be marshalled through
request/process/reply code.

But in this first step we know that the compaction code is running on
the server so we can forgo all the messaging and just call in to and out
of compaction.  The net calls just holds the ring consistency locks in
the server and call into the manifest to do the work, commiting the
changes when its done.

This is more careful about segno alloction and freeing.  Compaction
doesn't call the allocator directly.  It gets allocaitons from the
messages and returns them if it doesn't use them.  We actually now
free segnos as they're removed from the manifest.

With the server controlling compaction and can tear all the fiddly level
count watching code out of the manifest.  Item transactions can't care
about the level counts and the server always tries compaction after the
manifest is updated intead of having the manifest watch the level counts
and call compaction.

Now that the server owns the rings they should not be torn down as the
super is torn down, net does that now.  And we need to be more careful
to be sure that writes from dirtying and compaction are stable before
killing the super.

With all this in place moving to shared compaction involves adding the
messages and negotiating concurrent compactions in the manifest.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-24 14:02:18 -07:00
Zach Brown e09a216762 Support simpler ring entries
Add mkfs and print support for the simpler rings that the segment bitmap
allocator and manifest are now using.  Some other recent format header
updates come along for the ride.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown bd54995599 Add a simple native bitmap
Nothing fancy at all.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown f86ce74ffd Add BITS_PER_LONG define
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown a147239022 Remove dead block, btree, and buddy code
Remove the last bits of the dead code from the old btree design.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 2e2ee3b2f1 Print symlink items
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 77d0268cb2 Add printing xattrs
For now we only print the xattr names, not the values.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 13b2d9bb88 Remove find_xattr commands
We're no longer maintaining xattr backrefs.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 02993a2dd7 Update ino_path for the large cursor
Previously we could iterate over backref items with a small u64.  Now we
need a larger opaque buffer.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 16da3c182a Add printing link backref items
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown acda5a3bf1 Add support for free_segs in super
The allocator records the total number of free segments in the super
block.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 44f8551fb6 Print data items
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 52291b2c75 Update format for readdir_pos
We now track each parent dir's next readdir pos and the readdir pos of
each dirent.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 38c8a4901f Print orphan items
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown c4f2563cc1 Update tools to new segment item layout
The segment item struct used to have fiddly packed offsets and lengths.
Now it's just normal fields so we can work with them directly and get
rid of the native item indirection.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown e81c256a22 Remove the bitops helpers
We don't have any use for the bitops today, we'll resurrect this in
simpler form if it's needed again.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 34c62824e5 Use a treap walker to print segments
We were using a bitmap to record segments during manifest printing and
then walking that bitmap to print segments.  It's a little silly to have
a second data structure record the referenced segments when we could
just walk the manifest again to print the segments.

So refactor node printing into a treap walker that calls a function for
each node.  Then we can have functions that print the node data
structurs for each treap and then one that prints the segments that are
referenced by manifest nodes.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 26a4266964 Set manifest keys to precise segment keys
We had changed the manifest keys to fully cover the space around the
segments in the hopes that it'd let item reading easily find negative
cached regions around items.

But that makes compaction think that segments intersect with items when
they really don't.  We'd much rather avoid unnecessary compaction by
having the manifest entries precisely reflect the keys in the segment.

Item reading can do more work at run time to find the bounds of the key
space that are around the edges of the segments it works with.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown c2b47d84c1 Add next_seg_seq field to super
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 484b34057a Update mkfs and print for treap ring
Update mkfs and print now that the manifest and allocator are stored in
treaps in the ring.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:43 -07:00
Zach Brown 7c4bc528c6 Make sure manifests cover all keys
Make sure that the manifest entries for a given level fully
cover the possible key space.  This helps item reading describe
cached key ranges that extend around items.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:42 -07:00
Zach Brown c3b6dd0763 Describe ring log with index,nr
Update mkfs and print to describe the ring blocks with a starting index
and number of blocks instead of a head and tail index.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:42 -07:00
Zach Brown 19b674cb38 Print dirent and readdir items
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:42 -07:00
Zach Brown 7cd70ab2bb Don't double increment segno when printing
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:42 -07:00
Zach Brown 818e149643 Update mkfs and print for lsm writing
Adapt mkfs and print for the format changes made to support writing
segments.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:42 -07:00
Zach Brown eb4baa88f5 Print LSM structures
Print segments and their items instead of btree blocks.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:29 -07:00
Zach Brown c96b833a36 mkfs LSM segment and ring stuctures
Make a new file system by writing a root inode in a segment and storing
a manifest entry in the ring that references the segment.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 14:20:02 -07:00
Zach Brown 8b82aa7f18 Consistently initialize inode fields
Inode info struct initialization spread out over three places:
 - once for the memory of a slab obect
 - when reading an existing inode from items
 - when initializing a newly allocated inode

Over time field initializtion got out of sync with these rules.  This
makes it more clear which fields get initialized where.  In the inode
info struct we group fields by where there initialized.  We order the
fields by size and location in the inode struct.

Then we make sure that all the initialization sites have everything
covered.  Doing everything in consistent struct order makes it easier
to audit that we haven't missed anything.

What lead to this was realizing that we missed initializing the seqcount
when reading existing inodes.  It should have been initialized in the
slab object constructor.  The 'staging' boolean has the same problem.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:17:55 -07:00
Nic HenkeandZach Brown 5c54bdbf85 Change type for DATA_VERSION ioctl to __u64
For consistency and to keep upstream users (scout-utils, etc) from
needing to include different type headers, we'll change the type to
match the rest of the header.

Signed-off-by: Nic Henke <nic.henke@versity.com>
2017-04-18 14:07:23 -07:00
Zach Brown 37ba46213c Add suport for more xattr namespaces
Add support for more of the known xattr namespaces.  This helps
generic/062 in xfstests pass.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:06:29 -07:00
Zach Brown 2aa274b38b Add xattr iops for special files
xfstests generic/062 was failing because it was getting an unexpected
error code when trying to work with xattrs on special files.  Adding our
ops gives it the errnos it expects.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:06:29 -07:00
Zach Brown 78d15a019c Print inode nr and err on inode upate error
We're currently excessively freaking out if inode updates fail.  Let's
add a little more context to help us track down what goes wrong.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:06:26 -07:00
Zach Brown 2591e54fdc Make it easier to build scoutfs.ko
We were duplicating the make args a few times so make a little ARGS
variable.

Default to the /lib/modules/$(uname -r) installed kernel source if
SK_KSRC isn't set.

And only try a sparse build that can fail if we can execute the sparse
command.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:03:24 -07:00
Nic HenkeandZach Brown 9fc47dedf8 Add unlocked ioctls for directories.
The use of the Scout ioctls for inode-since and data-since on the root
directory is a rather helpful boost. This allows user code to start on
blank filesystems and monitor activity without needing to create files.

The existing ioctl code was already present, so wiring into the
directory file operations was all that needed to happen.

Signed-off-by: Nic Henke <nic.henke@versity.com>
Reviewed-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:03:24 -07:00
Zach Brown e61697a54e Add generic file and dir seek methods
Two more xfstests pass when we can seek in files and dirs.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 14:03:22 -07:00
Zach Brown efd95688d3 Add printf format checking to scoutfs msg funcs
scoutfs_msg() was missing the attribute to check printf formats and
arguments.

Signed-off-by: Zach Brown <zab@versity.com>
Reviewed-by: Mark Fasheh <mfasheh@versity.com>
2017-04-18 13:59:54 -07:00
Zach Brown cec3f9468a Further isolate rings and compaction
Each mount was still loading the manifest and allocator rings and
starting compaction, even if they were coordinating segment reads
and writes with the server.

This moves ring and compaction setup and teardown from on mount and
unmount to as the server starts up and shuts down.  Now only the server
has the rings resident and is running compaction.

We had to null some of the super info fields so that we can repeatedly
load and destroy the ring indices over the lifetime of a mount.

We also have to be careful not to call between item transactions and
compaction.   We'll restore this functionality with the server in the
future.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 5eefaf34f8 Server updates ring for level0 segment writes
Transaction commits currently directly modify the ring and super block
as segments are written.  As we introduce shared mounts only the server
can modify the ring and super blocks.

This adds network messages to let mounts write items in a level 0
segment while the server modifies the allocator and manifest.

The item transaction commit now sends a message to the server to get an
allocated segno for its new level0 segment and sends a manifest entry to
the server once the segment is written.  The request and reply handlers
for the functions are straight forward.  The processing paths are simple
wrappers around the allocation and update functions that transaction
writing used to call directly.

Now that the item transactions aren't updating the super sync can't
work with the super sequence numbers.

The server needs to make both allocations and manifest updates
persistent before it sends replies to the client.  We add the ability
for the server processing paths to queue and wait for commits of the
rings and super block.  We can hopefull get reasonable batching by using
a work struct for the commit.  We update the other processing path
callers that modify the rings to use the new commit mechanism.

We add a few segment and manifest functions to work with manifest
entries that describe segments.  This creats a bit of similar looking
code thorughout the segment and manifest code but we'll come back and
clean this up once we see what the final shared support looks like.

scoutfs_seg_alloc() now takes the segno from the caller for the segment
it's allocating and inserting into the cache.  Transaction commit uses
the segno it got from the server while compaction still allocates
locally.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 5487aee6a7 Read items with manifest entries from server
Item reading tries to directly walk the manifest to find segments to
read.  That doesn't work when only the server has read the ring and
loaded the manifest.

This adds a network message to ask the server for the manifest entries
that describe the segments that will be needed to read items.

Previously item reading would walk the manifest and build up native
manifest references in a list that it'd use to read.   To implement the
network message we add request sending, processing, and reply parsing
around those original functions.  Item reading now packs its key range
and sends it to the server.  The server walks the manifest and sends the
entries that intersect with the key range.  Then the reply function
builds up the native manifest references that item reading will use.

The net reply functions needed an argument so that the manifest reading
request could pass in the caller's list that the native manifest
references should be added to.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown b50de90196 Alloc inodes from pool from server
Inode allocation was always modifying the in-memory super block.  This
doesn't work when the server is solely responsible for modifying the
super blocks.  We add network messages to have mounts send a message to
the server to request inodes that they can use to satisfy allocation.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 453715a78d Only shutdown locks that were setup
Lock shutdown was crashing trying to deref a null linf on cleanup from
mont errors that happened before locks were setup.  Make sure lock
shutdown only tries to do work if the locks have been setup.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 45882f5a77 Add some ring tracing
Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 5e0e9ac12e Move to much simpler manifest/alloc storage
Using the treap to be able to incrementally read and write the manifest
and allocation storage from all nodes wasn't quite ready for prime time.
The biggest problem is that invalidating cached nodes which are the
target of native pointers, either for consistency or memory pressure, is
problematic.  This was getting in the way of adding shared support as
readers and writers try to use as much of their treap caches as they
can.  There were other serious problems that we'd run into eventually:
memory pressure from duplicate caching in native nodes and the page
cache, small IOs from reading a page at a time, the risk of
pathologically imbalanced treaps, and the ring being corrupted if the
migration balancing doesn't work (the model assumed you could always
dirty an individual node in a transaction, you have to dirty all the
parents in each new transaction).

Let's back off to a much simpler mechanism while we build the rest of
the system around it.  We can revisit aggressively optimizing this when
it's our worst problem.

We'll store the indexes that the manifest server needs in simple
preallocated rings with log entries.   The server has to read the index
in its entirety into a native rbtree before it can work on it.  We won't
access the physical ring from mounts anymore, they'll send messages to
the server.

The ring callers are now working with a pinned tree in memory so the
interface can be a bit simpler.  By storing the indexes in their own
rings the code and write path become a lot simper: we have an IO
submission path for each index instead of "dirtying" calls per index and
then a writing call.

All this is much more robust and much less likely to get in our way as
we stand up the rest of the system around it.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 86d3090982 Tighten lock range error handling
If lock_range returns an error then the caller won't unlock the range.
Make sure to unlock the range if we have it locked when we get errors
that we're going to return to the caller.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00
Zach Brown 104bbb06a9 Remove cached range when invalidating items
When invalidating items we need to remove the cached
range that covers the range of keys that we're removing so that
the removed items aren't then considered negative cached items.

Signed-off-by: Zach Brown <zab@versity.com>
2017-04-18 13:51:10 -07:00