From: Johannes Thumshirn <Johannes.Thumshirn@wdc.com>
To: "fdmanana@gmail.com" <fdmanana@gmail.com>,
Naohiro Aota <Naohiro.Aota@wdc.com>
Cc: linux-btrfs <linux-btrfs@vger.kernel.org>,
David Sterba <dsterba@suse.com>, "hare@suse.com" <hare@suse.com>,
linux-fsdevel <linux-fsdevel@vger.kernel.org>,
Jens Axboe <axboe@kernel.dk>,
"hch@infradead.org" <hch@infradead.org>,
"Darrick J. Wong" <darrick.wong@oracle.com>,
Josef Bacik <josef@toxicpanda.com>
Subject: Re: [PATCH v14 42/42] btrfs: reorder log node allocation
Date: Mon, 1 Feb 2021 15:54:16 +0000 [thread overview]
Message-ID: <SN4PR0401MB3598F91AA78520BE1946E9EC9BB69@SN4PR0401MB3598.namprd04.prod.outlook.com> (raw)
In-Reply-To: CAL3q7H7KYKgJgm2+C9WBW+F23tpdJukNZHQ8N-RxbyyC78B5xw@mail.gmail.com
On 01/02/2021 16:49, Filipe Manana wrote:
> On Wed, Jan 27, 2021 at 6:48 PM Naohiro Aota <naohiro.aota@wdc.com> wrote:
>>
>> This is the 3/3 patch to enable tree-log on ZONED mode.
>>
>> The allocation order of nodes of "fs_info->log_root_tree" and nodes of
>> "root->log_root" is not the same as the writing order of them. So, the
>> writing causes unaligned write errors.
>>
>> This patch reorders the allocation of them by delaying allocation of the
>> root node of "fs_info->log_root_tree," so that the node buffers can go out
>> sequentially to devices.
>>
>> Reviewed-by: Josef Bacik <josef@toxicpanda.com>
>> Signed-off-by: Johannes Thumshirn <johannes.thumshirn@wdc.com>
>> Signed-off-by: Naohiro Aota <naohiro.aota@wdc.com>
>> ---
>> fs/btrfs/disk-io.c | 7 -------
>> fs/btrfs/tree-log.c | 24 ++++++++++++++++++------
>> 2 files changed, 18 insertions(+), 13 deletions(-)
>>
>> diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c
>> index c3b5cfe4d928..d2b30716de84 100644
>> --- a/fs/btrfs/disk-io.c
>> +++ b/fs/btrfs/disk-io.c
>> @@ -1241,18 +1241,11 @@ int btrfs_init_log_root_tree(struct btrfs_trans_handle *trans,
>> struct btrfs_fs_info *fs_info)
>> {
>> struct btrfs_root *log_root;
>> - int ret;
>>
>> log_root = alloc_log_tree(trans, fs_info);
>> if (IS_ERR(log_root))
>> return PTR_ERR(log_root);
>>
>> - ret = btrfs_alloc_log_tree_node(trans, log_root);
>> - if (ret) {
>> - btrfs_put_root(log_root);
>> - return ret;
>> - }
>> -
>> WARN_ON(fs_info->log_root_tree);
>> fs_info->log_root_tree = log_root;
>> return 0;
>> diff --git a/fs/btrfs/tree-log.c b/fs/btrfs/tree-log.c
>> index 71a1c0b5bc26..d8315363dc1e 100644
>> --- a/fs/btrfs/tree-log.c
>> +++ b/fs/btrfs/tree-log.c
>> @@ -3159,6 +3159,16 @@ int btrfs_sync_log(struct btrfs_trans_handle *trans,
>> list_add_tail(&root_log_ctx.list, &log_root_tree->log_ctxs[index2]);
>> root_log_ctx.log_transid = log_root_tree->log_transid;
>>
>> + mutex_lock(&fs_info->tree_log_mutex);
>> + if (!log_root_tree->node) {
>> + ret = btrfs_alloc_log_tree_node(trans, log_root_tree);
>> + if (ret) {
>> + mutex_unlock(&fs_info->tree_log_mutex);
>> + goto out;
>> + }
>> + }
>> + mutex_unlock(&fs_info->tree_log_mutex);
>
> Hum, so this now has an impact for non-zoned mode.
>
> It reduces the parallelism between a previous transaction finishing
> its commit and an fsync started in the current transaction.
>
> A transaction commit releases fs_info->tree_log_mutex after it commits
> the super block.
>
> By taking that mutex here, we wait for the transaction commit to write
> the super blocks before we update the log root, start writeback of log
> tree nodes and wait for the writeback to complete.
>
> Before this change, we would do those 3 things before waiting for the
> previous transaction to commit the super blocks.
>
> So this undoes part of what was made by the following commit:
>
> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=47876f7ceffa0e6af7476e052b3c061f1f2c1d9f
>
> Which landed in 5.10. This patch and the rest of the patchset was
> based on pre 5.10 code - was this missed because of that?
> Or is there some other reason to have to do things this way for non-zoned mode?
>
> I think we should preserve the behaviour for non-zoned mode - i.e. any
> reason why not allocating log_root_tree->node at the top of start
> log_trans(), while under the protection of tree_root->log_mutex?
>
> My impression is that this, and the other patch with the subject
> "btrfs: serialize log transaction on ZONED mode", are out of sync with
> the changes in 5.10.
>
> Thanks, sorry for the late review.
Yes this code has been under development for quite some time and we probably
missed this upstream update.
We will be sending a fix ASAP.
next prev parent reply other threads:[~2021-02-01 15:55 UTC|newest]
Thread overview: 72+ messages / expand[flat|nested] mbox.gz Atom feed top
2021-01-26 2:24 [PATCH v14 00/42] btrfs: zoned block device support Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 01/42] block: add bio_add_zone_append_page Naohiro Aota
2021-01-26 16:08 ` Jens Axboe
2021-01-26 2:24 ` [PATCH v14 02/42] iomap: support REQ_OP_ZONE_APPEND Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 03/42] btrfs: defer loading zone info after opening trees Naohiro Aota
2021-01-30 22:09 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 04/42] btrfs: use regular SB location on emulated zoned mode Naohiro Aota
2021-01-30 22:28 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 05/42] btrfs: release path before calling into btrfs_load_block_group_zone_info Naohiro Aota
2021-01-30 23:21 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 06/42] btrfs: do not load fs_info->zoned from incompat flag Naohiro Aota
2021-01-30 23:40 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 07/42] btrfs: disallow fitrim in ZONED mode Naohiro Aota
2021-01-30 23:44 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 08/42] btrfs: allow zoned mode on non-zoned block devices Naohiro Aota
2021-01-31 1:17 ` Anand Jain
2021-02-01 11:06 ` Johannes Thumshirn
2021-02-02 1:49 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 09/42] btrfs: implement zoned chunk allocator Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 10/42] btrfs: verify device extent is aligned to zone Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 11/42] btrfs: load zone's allocation offset Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 12/42] btrfs: calculate allocation offset for conventional zones Naohiro Aota
2021-01-27 18:03 ` Josef Bacik
2021-02-03 5:19 ` Anand Jain
2021-02-03 6:10 ` Damien Le Moal
2021-02-03 6:56 ` Anand Jain
2021-02-03 7:10 ` Damien Le Moal
2021-01-26 2:24 ` [PATCH v14 13/42] btrfs: track unusable bytes for zones Naohiro Aota
2021-01-27 18:06 ` Josef Bacik
2021-01-26 2:24 ` [PATCH v14 14/42] btrfs: do sequential extent allocation in ZONED mode Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 15/42] btrfs: redirty released extent buffers " Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 16/42] btrfs: advance allocation pointer after tree log node Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 17/42] btrfs: enable to mount ZONED incompat flag Naohiro Aota
2021-01-31 12:21 ` Anand Jain
2021-01-26 2:24 ` [PATCH v14 18/42] btrfs: reset zones of unused block groups Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 19/42] btrfs: extract page adding function Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 20/42] btrfs: use bio_add_zone_append_page for zoned btrfs Naohiro Aota
2021-01-26 2:24 ` [PATCH v14 21/42] btrfs: handle REQ_OP_ZONE_APPEND as writing Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 22/42] btrfs: split ordered extent when bio is sent Naohiro Aota
2021-01-27 19:00 ` Josef Bacik
2021-01-26 2:25 ` [PATCH v14 23/42] btrfs: check if bio spans across an ordered extent Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 24/42] btrfs: extend btrfs_rmap_block for specifying a device Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 25/42] btrfs: cache if block-group is on a sequential zone Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 26/42] btrfs: save irq flags when looking up an ordered extent Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 27/42] btrfs: use ZONE_APPEND write for ZONED btrfs Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 28/42] btrfs: enable zone append writing for direct IO Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 29/42] btrfs: introduce dedicated data write path for ZONED mode Naohiro Aota
2021-02-02 15:00 ` David Sterba
2021-02-04 8:25 ` Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 30/42] btrfs: serialize meta IOs on " Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 31/42] btrfs: wait existing extents before truncating Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 32/42] btrfs: avoid async metadata checksum on ZONED mode Naohiro Aota
2021-02-02 14:54 ` David Sterba
2021-02-02 16:50 ` Johannes Thumshirn
2021-02-02 19:28 ` David Sterba
2021-01-26 2:25 ` [PATCH v14 33/42] btrfs: mark block groups to copy for device-replace Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 34/42] btrfs: implement cloning for ZONED device-replace Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 35/42] btrfs: implement copying " Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 36/42] btrfs: support dev-replace in ZONED mode Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 37/42] btrfs: enable relocation " Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 38/42] btrfs: relocate block group to repair IO failure in ZONED Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 39/42] btrfs: split alloc_log_tree() Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 40/42] btrfs: extend zoned allocator to use dedicated tree-log block group Naohiro Aota
2021-01-26 2:25 ` [PATCH v14 41/42] btrfs: serialize log transaction on ZONED mode Naohiro Aota
2021-01-27 19:01 ` Josef Bacik
2021-02-01 15:48 ` Filipe Manana
2021-01-26 2:25 ` [PATCH v14 42/42] btrfs: reorder log node allocation Naohiro Aota
2021-02-01 15:48 ` Filipe Manana
2021-02-01 15:54 ` Johannes Thumshirn [this message]
2021-01-29 7:56 ` [PATCH v14 00/42] btrfs: zoned block device support Johannes Thumshirn
2021-01-29 20:44 ` David Sterba
2021-01-30 11:30 ` Johannes Thumshirn
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=SN4PR0401MB3598F91AA78520BE1946E9EC9BB69@SN4PR0401MB3598.namprd04.prod.outlook.com \
--to=johannes.thumshirn@wdc.com \
--cc=Naohiro.Aota@wdc.com \
--cc=axboe@kernel.dk \
--cc=darrick.wong@oracle.com \
--cc=dsterba@suse.com \
--cc=fdmanana@gmail.com \
--cc=hare@suse.com \
--cc=hch@infradead.org \
--cc=josef@toxicpanda.com \
--cc=linux-btrfs@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).