From: Tejun Heo <tj@kernel.org> To: Michal Hocko <mhocko@kernel.org> Cc: Johannes Weiner <hannes@cmpxchg.org>, Petr Mladek <pmladek@suse.com>, cgroups@vger.kernel.org, Cyril Hrubis <chrubis@suse.cz>, linux-kernel@vger.kernel.org Subject: Re: [PATCH for-4.6-fixes] memcg: remove lru_add_drain_all() invocation from mem_cgroup_move_charge() Date: Wed, 20 Apr 2016 17:29:22 -0400 [thread overview] Message-ID: <20160420212922.GH4775@htj.duckdns.org> (raw) In-Reply-To: <20160417120747.GC21757@dhcp22.suse.cz> Hello, Michal. On Sun, Apr 17, 2016 at 08:07:48AM -0400, Michal Hocko wrote: > On Fri 15-04-16 15:17:19, Tejun Heo wrote: > > mem_cgroup_move_charge() invokes lru_add_drain_all() so that the pvec > > pages can be moved too. lru_add_drain_all() schedules and flushes > > work items on system_wq which depends on being able to create new > > kworkers to make forward progress. Since 1ed1328792ff ("sched, > > cgroup: replace signal_struct->group_rwsem with a global > > percpu_rwsem"), a new task can't be created while in the cgroup > > migration path and the described lru_add_drain_all() invocation can > > easily lead to a deadlock. > > > > Charge moving is best-effort and whether the pvec pages are migrated > > or not doesn't really matter. Don't call it during charge moving. > > Eventually, we want to move the actual charge moving outside the > > migration path. > > > > Signed-off-by: Tejun Heo <tj@kernel.org> > > Reported-by: Johannes Weiner <hannes@cmpxchg.org> > > I guess > Debugged-by: Petr Mladek <pmladek@suse.com> > Reported-by: Cyril Hrubis <chrubis@suse.cz> Yeah, definitely. Sorry about missing them. > > Suggested-by: Michal Hocko <mhocko@kernel.org> > > Fixes: 1ed1328792ff ("sched, cgroup: replace signal_struct->group_rwsem with a global percpu_rwsem") > > Cc: stable@vger.kernel.org > > --- > > Hello, > > > > So, this deadlock seems pretty easy to trigger. We'll make the charge > > moving asynchronous eventually but let's not hold off fixing an > > immediate problem. > > Although this looks rather straightforward and it fixes the immediate > problem I am little bit nervous about it. As already pointed out in > other email mem_cgroup_move_charge still depends on mmap_sem for > read and we might hit an even more subtle lockup if the current holder > of the mmap_sem for write depends on the task creation (e.g. some of the > direct reclaim path uses WQ which is really hard to rule out and I even > think that some shrinkers do this). > > I liked your proposal when mem_cgroup_move_charge would be called from a > context which doesn't hold the problematic rwsem much more. Would that > be too intrusive for the stable backport? Yeah, I'm working on the fix but let's plug this one first as it seems really easy to trigger. I got a couple off-list reports (in and outside fb) of this triggering. Thanks. -- tejun
WARNING: multiple messages have this Message-ID (diff)
From: Tejun Heo <tj-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org> To: Michal Hocko <mhocko-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org> Cc: Johannes Weiner <hannes-druUgvl0LCNAfugRpC6u6w@public.gmane.org>, Petr Mladek <pmladek-IBi9RG/b67k@public.gmane.org>, cgroups-u79uwXL29TY76Z2rM5mHXA@public.gmane.org, Cyril Hrubis <chrubis-AlSwsSmVLrQ@public.gmane.org>, linux-kernel-u79uwXL29TY76Z2rM5mHXA@public.gmane.org Subject: Re: [PATCH for-4.6-fixes] memcg: remove lru_add_drain_all() invocation from mem_cgroup_move_charge() Date: Wed, 20 Apr 2016 17:29:22 -0400 [thread overview] Message-ID: <20160420212922.GH4775@htj.duckdns.org> (raw) In-Reply-To: <20160417120747.GC21757-2MMpYkNvuYDjFM9bn6wA6Q@public.gmane.org> Hello, Michal. On Sun, Apr 17, 2016 at 08:07:48AM -0400, Michal Hocko wrote: > On Fri 15-04-16 15:17:19, Tejun Heo wrote: > > mem_cgroup_move_charge() invokes lru_add_drain_all() so that the pvec > > pages can be moved too. lru_add_drain_all() schedules and flushes > > work items on system_wq which depends on being able to create new > > kworkers to make forward progress. Since 1ed1328792ff ("sched, > > cgroup: replace signal_struct->group_rwsem with a global > > percpu_rwsem"), a new task can't be created while in the cgroup > > migration path and the described lru_add_drain_all() invocation can > > easily lead to a deadlock. > > > > Charge moving is best-effort and whether the pvec pages are migrated > > or not doesn't really matter. Don't call it during charge moving. > > Eventually, we want to move the actual charge moving outside the > > migration path. > > > > Signed-off-by: Tejun Heo <tj-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org> > > Reported-by: Johannes Weiner <hannes-druUgvl0LCNAfugRpC6u6w@public.gmane.org> > > I guess > Debugged-by: Petr Mladek <pmladek-IBi9RG/b67k@public.gmane.org> > Reported-by: Cyril Hrubis <chrubis-AlSwsSmVLrQ@public.gmane.org> Yeah, definitely. Sorry about missing them. > > Suggested-by: Michal Hocko <mhocko-DgEjT+Ai2ygdnm+yROfE0A@public.gmane.org> > > Fixes: 1ed1328792ff ("sched, cgroup: replace signal_struct->group_rwsem with a global percpu_rwsem") > > Cc: stable-u79uwXL29TY76Z2rM5mHXA@public.gmane.org > > --- > > Hello, > > > > So, this deadlock seems pretty easy to trigger. We'll make the charge > > moving asynchronous eventually but let's not hold off fixing an > > immediate problem. > > Although this looks rather straightforward and it fixes the immediate > problem I am little bit nervous about it. As already pointed out in > other email mem_cgroup_move_charge still depends on mmap_sem for > read and we might hit an even more subtle lockup if the current holder > of the mmap_sem for write depends on the task creation (e.g. some of the > direct reclaim path uses WQ which is really hard to rule out and I even > think that some shrinkers do this). > > I liked your proposal when mem_cgroup_move_charge would be called from a > context which doesn't hold the problematic rwsem much more. Would that > be too intrusive for the stable backport? Yeah, I'm working on the fix but let's plug this one first as it seems really easy to trigger. I got a couple off-list reports (in and outside fb) of this triggering. Thanks. -- tejun
next prev parent reply other threads:[~2016-04-20 21:29 UTC|newest] Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top 2016-04-13 9:42 [BUG] cgroup/workques/fork: deadlock when moving cgroups Petr Mladek 2016-04-13 18:33 ` Tejun Heo 2016-04-13 18:33 ` Tejun Heo 2016-04-13 18:57 ` Tejun Heo 2016-04-13 18:57 ` Tejun Heo 2016-04-13 19:23 ` Michal Hocko 2016-04-13 19:23 ` Michal Hocko 2016-04-13 19:28 ` Michal Hocko 2016-04-13 19:28 ` Michal Hocko 2016-04-13 19:37 ` Tejun Heo 2016-04-13 19:48 ` Michal Hocko 2016-04-14 7:06 ` Michal Hocko 2016-04-14 7:06 ` Michal Hocko 2016-04-14 15:32 ` Tejun Heo 2016-04-14 15:32 ` Tejun Heo 2016-04-14 17:50 ` Johannes Weiner 2016-04-15 7:06 ` Michal Hocko 2016-04-15 14:38 ` Tejun Heo 2016-04-15 14:38 ` Tejun Heo 2016-04-15 15:08 ` Michal Hocko 2016-04-15 15:08 ` Michal Hocko 2016-04-15 15:25 ` Tejun Heo 2016-04-15 15:25 ` Tejun Heo 2016-04-17 12:00 ` Michal Hocko 2016-04-17 12:00 ` Michal Hocko 2016-04-18 14:40 ` Petr Mladek 2016-04-18 14:40 ` Petr Mladek 2016-04-19 14:01 ` Michal Hocko 2016-04-19 14:01 ` Michal Hocko 2016-04-19 15:39 ` Petr Mladek 2016-04-15 19:17 ` [PATCH for-4.6-fixes] memcg: remove lru_add_drain_all() invocation from mem_cgroup_move_charge() Tejun Heo 2016-04-17 12:07 ` Michal Hocko 2016-04-17 12:07 ` Michal Hocko 2016-04-20 21:29 ` Tejun Heo [this message] 2016-04-20 21:29 ` Tejun Heo 2016-04-21 3:27 ` Michal Hocko 2016-04-21 3:27 ` Michal Hocko 2016-04-21 15:00 ` Petr Mladek 2016-04-21 15:00 ` Petr Mladek 2016-04-21 15:51 ` Tejun Heo 2016-04-21 23:06 ` [PATCH 1/2] cgroup, cpuset: replace cpuset_post_attach_flush() with cgroup_subsys->post_attach callback Tejun Heo 2016-04-21 23:06 ` Tejun Heo 2016-04-21 23:09 ` [PATCH 2/2] memcg: relocate charge moving from ->attach to ->post_attach Tejun Heo 2016-04-21 23:09 ` Tejun Heo 2016-04-22 13:57 ` Petr Mladek 2016-04-22 13:57 ` Petr Mladek 2016-04-25 8:25 ` Michal Hocko 2016-04-25 8:25 ` Michal Hocko 2016-04-25 19:42 ` Tejun Heo 2016-04-25 19:42 ` Tejun Heo 2016-04-25 19:44 ` Tejun Heo 2016-04-25 19:44 ` Tejun Heo 2016-04-21 23:11 ` [PATCH 1/2] cgroup, cpuset: replace cpuset_post_attach_flush() with cgroup_subsys->post_attach callback Tejun Heo 2016-04-21 23:11 ` Tejun Heo 2016-04-21 15:56 ` [PATCH for-4.6-fixes] memcg: remove lru_add_drain_all() invocation from mem_cgroup_move_charge() Tejun Heo 2016-04-21 15:56 ` Tejun Heo
Reply instructions: You may reply publicly to this message via plain-text email using any one of the following methods: * Save the following mbox file, import it into your mail client, and reply-to-all from there: mbox Avoid top-posting and favor interleaved quoting: https://en.wikipedia.org/wiki/Posting_style#Interleaved_style * Reply using the --to, --cc, and --in-reply-to switches of git-send-email(1): git send-email \ --in-reply-to=20160420212922.GH4775@htj.duckdns.org \ --to=tj@kernel.org \ --cc=cgroups@vger.kernel.org \ --cc=chrubis@suse.cz \ --cc=hannes@cmpxchg.org \ --cc=linux-kernel@vger.kernel.org \ --cc=mhocko@kernel.org \ --cc=pmladek@suse.com \ /path/to/YOUR_REPLY https://kernel.org/pub/software/scm/git/docs/git-send-email.html * If your mail client supports setting the In-Reply-To header via mailto: links, try the mailto: linkBe sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes, see mirroring instructions on how to clone and mirror all data and code used by this external index.