From: Minchan Kim <minchan@kernel.org>
To: Michal Hocko <mhocko@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
LKML <linux-kernel@vger.kernel.org>,
linux-mm <linux-mm@kvack.org>,
Johannes Weiner <hannes@cmpxchg.org>,
Tim Murray <timmurray@google.com>,
Joel Fernandes <joel@joelfernandes.org>,
Suren Baghdasaryan <surenb@google.com>,
Daniel Colascione <dancol@google.com>,
Shakeel Butt <shakeelb@google.com>,
Sonny Rao <sonnyrao@google.com>,
Brian Geffon <bgeffon@google.com>,
linux-api@vger.kernel.org
Subject: Re: [RFC 5/7] mm: introduce external memory hinting API
Date: Tue, 21 May 2019 19:32:23 +0900 [thread overview]
Message-ID: <20190521103223.GD219653@google.com> (raw)
In-Reply-To: <20190521061743.GC32329@dhcp22.suse.cz>
On Tue, May 21, 2019 at 08:17:43AM +0200, Michal Hocko wrote:
> On Tue 21-05-19 11:41:07, Minchan Kim wrote:
> > On Mon, May 20, 2019 at 11:18:29AM +0200, Michal Hocko wrote:
> > > [Cc linux-api]
> > >
> > > On Mon 20-05-19 12:52:52, Minchan Kim wrote:
> > > > There is some usecase that centralized userspace daemon want to give
> > > > a memory hint like MADV_[COOL|COLD] to other process. Android's
> > > > ActivityManagerService is one of them.
> > > >
> > > > It's similar in spirit to madvise(MADV_WONTNEED), but the information
> > > > required to make the reclaim decision is not known to the app. Instead,
> > > > it is known to the centralized userspace daemon(ActivityManagerService),
> > > > and that daemon must be able to initiate reclaim on its own without
> > > > any app involvement.
> > >
> > > Could you expand some more about how this all works? How does the
> > > centralized daemon track respective ranges? How does it synchronize
> > > against parallel modification of the address space etc.
> >
> > Currently, we don't track each address ranges because we have two
> > policies at this moment:
> >
> > deactive file pages and reclaim anonymous pages of the app.
> >
> > Since the daemon has a ability to let background apps resume(IOW, process
> > will be run by the daemon) and both hints are non-disruptive stabilty point
> > of view, we are okay for the race.
>
> Fair enough but the API should consider future usecases where this might
> be a problem. So we should really think about those potential scenarios
> now. If we are ok with that, fine, but then we should be explicit and
> document it that way. Essentially say that any sort of synchronization
> is supposed to be done by monitor. This will make the API less usable
> but maybe that is enough.
Okay, I will add more about that in the description.
>
> > > > To solve the issue, this patch introduces new syscall process_madvise(2)
> > > > which works based on pidfd so it could give a hint to the exeternal
> > > > process.
> > > >
> > > > int process_madvise(int pidfd, void *addr, size_t length, int advise);
> > >
> > > OK, this makes some sense from the API point of view. When we have
> > > discussed that at LSFMM I was contemplating about something like that
> > > except the fd would be a VMA fd rather than the process. We could extend
> > > and reuse /proc/<pid>/map_files interface which doesn't support the
> > > anonymous memory right now.
> > >
> > > I am not saying this would be a better interface but I wanted to mention
> > > it here for a further discussion. One slight advantage would be that
> > > you know the exact object that you are operating on because you have a
> > > fd for the VMA and we would have a more straightforward way to reject
> > > operation if the underlying object has changed (e.g. unmapped and reused
> > > for a different mapping).
> >
> > I agree your point. If I didn't miss something, such kinds of vma level
> > modify notification doesn't work even file mapped vma at this moment.
> > For anonymous vma, I think we could use userfaultfd, pontentially.
> > It would be great if someone want to do with disruptive hints like
> > MADV_DONTNEED.
> >
> > I'd like to see it further enhancement after landing address range based
> > operation via limiting hints process_madvise supports to non-disruptive
> > only(e.g., MADV_[COOL|COLD]) so that we could catch up the usercase/workload
> > when someone want to extend the API.
>
> So do you think we want both interfaces (process_madvise and madvisefd)?
What I have in mind is to extend process_madvise later like this
struct pr_madvise_param {
int size; /* the size of this structure */
union {
const struct iovec __user *vec; /* address range array */
int fd; /* supported from 6.0 */
}
}
with introducing new hint Or-able PR_MADV_RANGE_FD, so that process_madvise
can go with fd instead of address range.
next prev parent reply other threads:[~2019-05-21 10:32 UTC|newest]
Thread overview: 138+ messages / expand[flat|nested] mbox.gz Atom feed top
2019-05-20 3:52 [RFC 0/7] introduce memory hinting API for external process Minchan Kim
2019-05-20 3:52 ` [RFC 1/7] mm: introduce MADV_COOL Minchan Kim
2019-05-20 8:16 ` Michal Hocko
2019-05-20 8:19 ` Michal Hocko
2019-05-20 15:08 ` Suren Baghdasaryan
2019-05-20 22:55 ` Minchan Kim
2019-05-20 22:54 ` Minchan Kim
2019-05-21 6:04 ` Michal Hocko
2019-05-21 9:11 ` Minchan Kim
2019-05-21 10:05 ` Michal Hocko
2019-05-28 8:53 ` Hillf Danton
2019-05-28 10:58 ` Minchan Kim
2019-05-20 3:52 ` [RFC 2/7] mm: change PAGEREF_RECLAIM_CLEAN with PAGE_REFRECLAIM Minchan Kim
2019-05-20 16:50 ` Johannes Weiner
2019-05-20 22:57 ` Minchan Kim
2019-05-20 3:52 ` [RFC 3/7] mm: introduce MADV_COLD Minchan Kim
2019-05-20 8:27 ` Michal Hocko
2019-05-20 23:00 ` Minchan Kim
2019-05-21 6:08 ` Michal Hocko
2019-05-21 9:13 ` Minchan Kim
2019-05-28 14:54 ` Hillf Danton
2019-05-30 0:45 ` Minchan Kim
2019-05-20 3:52 ` [RFC 4/7] mm: factor out madvise's core functionality Minchan Kim
2019-05-20 14:26 ` Oleksandr Natalenko
2019-05-21 1:26 ` Minchan Kim
2019-05-21 6:36 ` Oleksandr Natalenko
2019-05-21 6:50 ` Michal Hocko
2019-05-21 7:06 ` Oleksandr Natalenko
2019-05-21 10:52 ` Minchan Kim
2019-05-21 11:00 ` Michal Hocko
2019-05-21 11:24 ` Minchan Kim
2019-05-21 11:32 ` Michal Hocko
2019-05-21 10:49 ` Minchan Kim
2019-05-21 10:55 ` Michal Hocko
2019-05-20 3:52 ` [RFC 5/7] mm: introduce external memory hinting API Minchan Kim
2019-05-20 9:18 ` Michal Hocko
2019-05-21 2:41 ` Minchan Kim
2019-05-21 6:17 ` Michal Hocko
2019-05-21 10:32 ` Minchan Kim [this message]
2019-05-21 9:01 ` Christian Brauner
2019-05-21 11:35 ` Minchan Kim
2019-05-21 11:51 ` Christian Brauner
2019-05-21 15:31 ` Oleg Nesterov
2019-05-27 7:43 ` Minchan Kim
2019-05-27 15:12 ` Oleg Nesterov
2019-05-27 23:33 ` Minchan Kim
2019-05-28 7:23 ` Michal Hocko
2019-05-29 3:41 ` Hillf Danton
2019-05-30 0:38 ` Minchan Kim
2019-05-20 3:52 ` [RFC 6/7] mm: extend process_madvise syscall to support vector arrary Minchan Kim
2019-05-20 9:22 ` Michal Hocko
2019-05-21 2:48 ` Minchan Kim
2019-05-21 6:24 ` Michal Hocko
2019-05-21 10:26 ` Minchan Kim
2019-05-21 10:37 ` Michal Hocko
2019-05-27 7:49 ` Minchan Kim
2019-05-29 10:08 ` Daniel Colascione
2019-05-29 10:33 ` Michal Hocko
2019-05-30 2:17 ` Minchan Kim
2019-05-30 6:57 ` Michal Hocko
2019-05-30 8:02 ` Minchan Kim
2019-05-30 16:19 ` Daniel Colascione
2019-05-30 18:47 ` Michal Hocko
2019-05-29 4:14 ` Hillf Danton
2019-05-30 0:35 ` Minchan Kim
2019-05-20 3:52 ` [RFC 7/7] mm: madvise support MADV_ANONYMOUS_FILTER and MADV_FILE_FILTER Minchan Kim
2019-05-20 9:28 ` Michal Hocko
2019-05-21 2:55 ` Minchan Kim
2019-05-21 6:26 ` Michal Hocko
2019-05-27 7:58 ` Minchan Kim
2019-05-27 12:44 ` Michal Hocko
2019-05-28 3:26 ` Minchan Kim
2019-05-28 6:29 ` Michal Hocko
2019-05-28 8:13 ` Minchan Kim
2019-05-28 8:31 ` Daniel Colascione
2019-05-28 8:49 ` Minchan Kim
2019-05-28 9:08 ` Michal Hocko
2019-05-28 9:39 ` Daniel Colascione
2019-05-28 10:33 ` Michal Hocko
2019-05-28 11:21 ` Daniel Colascione
2019-05-28 11:49 ` Michal Hocko
2019-05-28 12:11 ` Daniel Colascione
2019-05-28 12:32 ` Michal Hocko
2019-05-28 10:32 ` Minchan Kim
2019-05-28 10:41 ` Michal Hocko
2019-05-28 11:12 ` Minchan Kim
2019-05-28 11:28 ` Michal Hocko
2019-05-28 11:42 ` Daniel Colascione
2019-05-28 11:56 ` Michal Hocko
2019-05-28 12:18 ` Daniel Colascione
2019-05-28 12:38 ` Michal Hocko
2019-05-28 12:10 ` Minchan Kim
2019-05-28 11:44 ` Minchan Kim
2019-05-28 11:51 ` Daniel Colascione
2019-05-28 12:06 ` Michal Hocko
2019-05-28 12:22 ` Minchan Kim
2019-05-28 11:28 ` Daniel Colascione
2019-05-21 15:33 ` Johannes Weiner
2019-05-22 1:50 ` Minchan Kim
2019-05-29 4:36 ` Hillf Danton
2019-05-30 1:00 ` Minchan Kim
2019-05-20 6:37 ` [RFC 0/7] introduce memory hinting API for external process Anshuman Khandual
2019-05-20 16:59 ` Tim Murray
2019-05-21 2:55 ` Anshuman Khandual
2019-05-21 5:14 ` Minchan Kim
2019-05-21 10:34 ` Michal Hocko
2019-05-28 10:50 ` Anshuman Khandual
2019-05-21 12:56 ` Shakeel Butt
2019-05-22 4:15 ` Brian Geffon
2019-05-22 4:23 ` Brian Geffon
2019-05-20 9:28 ` Michal Hocko
2019-05-20 14:42 ` Oleksandr Natalenko
2019-05-21 2:56 ` Minchan Kim
2019-05-20 16:46 ` Johannes Weiner
2019-05-21 4:39 ` Minchan Kim
2019-05-21 6:32 ` Michal Hocko
2019-05-21 1:44 ` Matthew Wilcox
2019-05-21 5:01 ` Minchan Kim
2019-05-21 6:34 ` Michal Hocko
2019-05-21 8:42 ` Christian Brauner
2019-05-21 11:05 ` Minchan Kim
2019-05-21 11:30 ` Christian Brauner
2019-05-21 11:39 ` Christian Brauner
2019-05-22 5:11 ` Daniel Colascione
2019-05-22 8:22 ` Christian Brauner
2019-05-22 13:16 ` Daniel Colascione
2019-05-22 14:52 ` Christian Brauner
2019-05-22 15:17 ` Daniel Colascione
2019-05-22 15:48 ` Christian Brauner
2019-05-22 15:57 ` Daniel Colascione
2019-05-22 16:01 ` Christian Brauner
2019-05-22 16:01 ` Daniel Colascione
2019-05-23 13:07 ` Minchan Kim
2019-05-27 8:06 ` Minchan Kim
2019-05-21 11:41 ` Minchan Kim
2019-05-21 12:04 ` Christian Brauner
2019-05-21 12:15 ` Oleksandr Natalenko
2019-05-21 12:53 ` Shakeel Butt
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20190521103223.GD219653@google.com \
--to=minchan@kernel.org \
--cc=akpm@linux-foundation.org \
--cc=bgeffon@google.com \
--cc=dancol@google.com \
--cc=hannes@cmpxchg.org \
--cc=joel@joelfernandes.org \
--cc=linux-api@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mhocko@kernel.org \
--cc=shakeelb@google.com \
--cc=sonnyrao@google.com \
--cc=surenb@google.com \
--cc=timmurray@google.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).