scoutfs: add precise transation item reservations

We had a simple mechanism for ensuring that transaction didn't create
more items than would fit in a single written segment.  We calculated
the most dirty items that a holder could generate and assumed that all
holders dirtied that much.

This had two big problems.

The first was that it wasn't accounting for nested holds.
write_begin/end calls the generic inode dirtying path whild holding a
transaction.  This ended up deadlocking as the dirty inode waited to be
able to write while its trans held back in write_begin prevented
writeout.

The second was that the worst case (full size xattr) item dirtying is
enormous and meaningfully restricts concurrent transaction holders.
With no currently dirty items you can have less than 16 full size xattr
writes.  This concurrency limit only gets worse as the transaction fills
up with dirty items.

This fixes those problems.  It adds precise accounting of the dirty
items that can be created while a transaction is held.  These
reservations are tracked in journal_info so that they can be used by
nested holds.  The precision allows much greater concurrency as
something like a create will try to reserve a few hundreds bytes instead
of 64k.  Normal sized xattr operations won't try to reserve the largest
possible space.

We add some feedback from the item cache to the transaction to issue
warnings if a holder dirties more items than it reserved.

Now that we have precise item/key/value counts (segment space
consumption is a function of all three :/) we can't have a single atomic
track transaction holders.  We add a long-overdue trans_info and put a
proper lock and fields there and much more clearly track transaction
serialization amongst the holders and writer.

Signed-off-by: Zach Brown <zab@versity.com>
This commit is contained in:
Zach Brown
2017-05-23 12:15:13 -07:00
parent 297b859577
commit b7bbad1fba
13 changed files with 452 additions and 111 deletions
+7 -2
View File
@@ -574,9 +574,11 @@ void scoutfs_update_inode_item(struct inode *inode)
void scoutfs_dirty_inode(struct inode *inode, int flags)
{
struct super_block *sb = inode->i_sb;
DECLARE_ITEM_COUNT(cnt);
int ret;
ret = scoutfs_hold_trans(sb);
scoutfs_count_dirty_inode(&cnt);
ret = scoutfs_hold_trans(sb, &cnt);
if (ret == 0) {
ret = scoutfs_dirty_inode_item(inode);
if (ret == 0)
@@ -777,12 +779,15 @@ static int remove_orphan_item(struct super_block *sb, u64 ino)
static int __delete_inode(struct super_block *sb, struct scoutfs_key_buf *key,
u64 ino, umode_t mode)
{
DECLARE_ITEM_COUNT(cnt);
bool release = false;
int ret;
trace_delete_inode(sb, ino, mode);
ret = scoutfs_hold_trans(sb);
/* XXX this is obviously not done yet :) */
scoutfs_count_dirty_inode(&cnt);
ret = scoutfs_hold_trans(sb, &cnt);
if (ret)
goto out;
release = true;