From: Shakeel Butt <shakeelb@google.com>
To: Johannes Weiner <hannes@cmpxchg.org>
Cc: "Alex Shi" <alex.shi@linux.alibaba.com>,
"Andrew Morton" <akpm@linux-foundation.org>,
Cgroups <cgroups@vger.kernel.org>,
LKML <linux-kernel@vger.kernel.org>,
"Linux MM" <linux-mm@kvack.org>,
"Mel Gorman" <mgorman@techsingularity.net>,
"Tejun Heo" <tj@kernel.org>, "Hugh Dickins" <hughd@google.com>,
"Konstantin Khlebnikov" <khlebnikov@yandex-team.ru>,
"Daniel Jordan" <daniel.m.jordan@oracle.com>,
"Yang Shi" <yang.shi@linux.alibaba.com>,
"Matthew Wilcox" <willy@infradead.org>,
"Michal Hocko" <mhocko@kernel.org>,
"Vladimir Davydov" <vdavydov.dev@gmail.com>,
"Roman Gushchin" <guro@fb.com>,
"Chris Down" <chris@chrisdown.name>,
"Thomas Gleixner" <tglx@linutronix.de>,
"Vlastimil Babka" <vbabka@suse.cz>, "Qian Cai" <cai@lca.pw>,
"Andrey Ryabinin" <aryabinin@virtuozzo.com>,
"Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>,
"Jérôme Glisse" <jglisse@redhat.com>,
"Andrea Arcangeli" <aarcange@redhat.com>,
"David Rientjes" <rientjes@google.com>,
"Aneesh Kumar K.V" <aneesh.kumar@linux.ibm.com>,
swkhack <swkhack@gmail.com>,
"Potyra, Stefan" <Stefan.Potyra@elektrobit.com>,
"Mike Rapoport" <rppt@linux.vnet.ibm.com>,
"Stephen Rothwell" <sfr@canb.auug.org.au>,
"Colin Ian King" <colin.king@canonical.com>,
"Jason Gunthorpe" <jgg@ziepe.ca>,
"Mauro Carvalho Chehab" <mchehab+samsung@kernel.org>,
"Peng Fan" <peng.fan@nxp.com>,
"Nikolay Borisov" <nborisov@suse.com>,
"Ira Weiny" <ira.weiny@intel.com>,
"Kirill Tkhai" <ktkhai@virtuozzo.com>,
"Yafang Shao" <laoar.shao@gmail.com>,
"Wei Yang" <richard.weiyang@gmail.com>
Subject: Re: [PATCH v8 03/10] mm/lru: replace pgdat lru_lock with lruvec lock
Date: Thu, 16 Apr 2020 10:47:00 -0700 [thread overview]
Message-ID: <CALvZod4bdmkd_YG=96O8+zCSCFNpsBQiN+3Cq+6oD7jn3GTYog@mail.gmail.com> (raw)
In-Reply-To: <20200416152830.GA195132@cmpxchg.org>
Hi Johannes & Alex,
On Thu, Apr 16, 2020 at 8:28 AM Johannes Weiner <hannes@cmpxchg.org> wrote:
>
> Hi Alex,
>
> On Thu, Apr 16, 2020 at 04:01:20PM +0800, Alex Shi wrote:
> >
> >
> > 在 2020/4/15 下午9:42, Alex Shi 写道:
> > > Hi Johannes,
> > >
> > > Thanks a lot for point out!
> > >
> > > Charging in __read_swap_cache_async would ask for 3 layers function arguments
> > > pass, that would be a bit ugly. Compare to this, could we move out the
> > > lru_cache add after commit_charge, like ksm copied pages?
> > >
> > > That give a bit extra non lru list time, but the page just only be used only
> > > after add_anon_rmap setting. Could it cause troubles?
> >
> > Hi Johannes & Andrew,
> >
> > Doing lru_cache_add_anon during swapin_readahead can give a very short timing
> > for possible page reclaiming for these few pages.
> >
> > If we delay these few pages lru adding till after the vm_fault target page
> > get memcg charging(mem_cgroup_commit_charge) and activate, we could skip the
> > mem_cgroup_try_charge/commit_charge/cancel_charge process in __read_swap_cache_async().
> > But the cost is maximum SWAP_RA_ORDER_CEILING number pages on each cpu miss
> > page reclaiming in a short time. On the other hand, save the target vm_fault
> > page from reclaiming before activate it during that time.
>
> The readahead pages surrounding the faulting page might never get
> accessed and pile up to large amounts. Users can also trigger
> non-faulting readahead with MADV_WILLNEED.
>
> So unfortunately, I don't see a way to keep these pages off the
> LRU. They do need to be reclaimable, or they become a DoS vector.
>
> I'm currently preparing a small patch series to make swap ownership
> tracking an integral part of memcg and change the swapin charging
> sequence, then you don't have to worry about it. This will also
> unblock Joonsoo's "workingset protection/detection on the anonymous
> LRU list" patch series, since he is blocked on the same problem - he
> needs the correct LRU available at swapin time to process refaults
> correctly. Both of your patch series are already pretty large, they
> shouldn't need to also deal with that.
I think this would be a very good cleanup and will make the code much
more readable. I totally agree to keep this separate from the other
work. Please do CC me the series once it's ready.
Now regarding the per-memcg LRU locks, Alex, did you get the chance to
try the workload Hugh has provided? I was planning of posting Hugh's
patch series but Hugh advised me to wait for your & Johannes's
response since you both have already invested a lot of time in your
series and I do want to see how Johannes's TestClearPageLRU() idea
will look like, so, I will hold off for now.
thanks,
Shakeel
next prev parent reply other threads:[~2020-04-16 17:47 UTC|newest]
Thread overview: 29+ messages / expand[flat|nested] mbox.gz Atom feed top
2020-01-16 3:04 [PATCH v8 00/10] per lruvec lru_lock for memcg Alex Shi
2020-01-16 3:05 ` [PATCH v8 01/10] mm/vmscan: remove unnecessary lruvec adding Alex Shi
2020-01-16 3:05 ` [PATCH v8 02/10] mm/memcg: fold lock_page_lru into commit_charge Alex Shi
2020-01-16 3:05 ` [PATCH v8 03/10] mm/lru: replace pgdat lru_lock with lruvec lock Alex Shi
2020-01-16 21:52 ` Johannes Weiner
2020-01-19 11:32 ` Alex Shi
2020-01-20 12:58 ` Alex Shi
2020-01-21 16:00 ` Johannes Weiner
2020-01-22 12:01 ` Alex Shi
2020-01-22 18:31 ` Johannes Weiner
2020-04-13 10:48 ` Alex Shi
2020-04-13 18:07 ` Johannes Weiner
2020-04-14 4:52 ` Alex Shi
2020-04-14 16:31 ` Johannes Weiner
2020-04-15 13:42 ` Alex Shi
2020-04-16 8:01 ` Alex Shi
2020-04-16 15:28 ` Johannes Weiner
2020-04-16 17:47 ` Shakeel Butt [this message]
2020-04-17 13:18 ` Alex Shi
2020-04-17 14:39 ` Alex Shi
2020-04-14 8:19 ` Alex Shi
2020-04-14 16:36 ` Johannes Weiner
2020-01-16 3:05 ` [PATCH v8 04/10] mm/lru: introduce the relock_page_lruvec function Alex Shi
2020-01-16 3:05 ` [PATCH v8 05/10] mm/mlock: optimize munlock_pagevec by relocking Alex Shi
2020-01-16 3:05 ` [PATCH v8 06/10] mm/swap: only change the lru_lock iff page's lruvec is different Alex Shi
2020-01-16 3:05 ` [PATCH v8 07/10] mm/pgdat: remove pgdat lru_lock Alex Shi
2020-01-16 3:05 ` [PATCH v8 08/10] mm/lru: revise the comments of lru_lock Alex Shi
2020-01-16 3:05 ` [PATCH v8 09/10] mm/lru: add debug checking for page memcg moving Alex Shi
2020-01-16 3:05 ` [PATCH v8 10/10] mm/memcg: add debug checking in lock_page_memcg Alex Shi
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to='CALvZod4bdmkd_YG=96O8+zCSCFNpsBQiN+3Cq+6oD7jn3GTYog@mail.gmail.com' \
--to=shakeelb@google.com \
--cc=Stefan.Potyra@elektrobit.com \
--cc=aarcange@redhat.com \
--cc=akpm@linux-foundation.org \
--cc=alex.shi@linux.alibaba.com \
--cc=aneesh.kumar@linux.ibm.com \
--cc=aryabinin@virtuozzo.com \
--cc=cai@lca.pw \
--cc=cgroups@vger.kernel.org \
--cc=chris@chrisdown.name \
--cc=colin.king@canonical.com \
--cc=daniel.m.jordan@oracle.com \
--cc=guro@fb.com \
--cc=hannes@cmpxchg.org \
--cc=hughd@google.com \
--cc=ira.weiny@intel.com \
--cc=jgg@ziepe.ca \
--cc=jglisse@redhat.com \
--cc=khlebnikov@yandex-team.ru \
--cc=kirill.shutemov@linux.intel.com \
--cc=ktkhai@virtuozzo.com \
--cc=laoar.shao@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mchehab+samsung@kernel.org \
--cc=mgorman@techsingularity.net \
--cc=mhocko@kernel.org \
--cc=nborisov@suse.com \
--cc=peng.fan@nxp.com \
--cc=richard.weiyang@gmail.com \
--cc=rientjes@google.com \
--cc=rppt@linux.vnet.ibm.com \
--cc=sfr@canb.auug.org.au \
--cc=swkhack@gmail.com \
--cc=tglx@linutronix.de \
--cc=tj@kernel.org \
--cc=vbabka@suse.cz \
--cc=vdavydov.dev@gmail.com \
--cc=willy@infradead.org \
--cc=yang.shi@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).