From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 4F93BC0502C for ; Wed, 31 Aug 2022 15:28:56 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S231593AbiHaP2z (ORCPT ); Wed, 31 Aug 2022 11:28:55 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:46192 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S230078AbiHaP2y (ORCPT ); Wed, 31 Aug 2022 11:28:54 -0400 Received: from mail-yw1-x112b.google.com (mail-yw1-x112b.google.com [IPv6:2607:f8b0:4864:20::112b]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id BB8FFD7D3D for ; Wed, 31 Aug 2022 08:28:52 -0700 (PDT) Received: by mail-yw1-x112b.google.com with SMTP id 00721157ae682-333a4a5d495so310099157b3.10 for ; Wed, 31 Aug 2022 08:28:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc; bh=WG3L5AqBV6fri0E+FNG2tsbDp/uFyzzRTMikwR7NSY8=; b=K4mZtzeFsKxJKOsvfz3u8buWcDbdHB0bPmoDat2amZVxf6i7fSAQFaeyCHfBhOW0gn TYiPTPwinxH8yKQGmCabK7zsAg7t70798MwIPLUmtbWdWtkPoH8nkkd0L6+Cf/71NitT pIpTm135oSjNuBnp7lYQC08Y98tRGNR+0DcT0LYO7fxV5Ng1p7CsbsO1uZH/I63XPhjj LoKETIXPqOU+NyCB1s85tiwuh2e//ou8VKmFXVH3pOn419Mn7E1W55VRw4JlXGoTp7nX lwPAdy+2tjWZHnwP0BxNNwETNl+fZjmT1bVCkdUM4qWebnjcYikmadptjs0201jXoXGN /Wug== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-message-state:from:to:cc; bh=WG3L5AqBV6fri0E+FNG2tsbDp/uFyzzRTMikwR7NSY8=; b=6sG0JLE0bQFKydyt2QHE/dBK+kxjkKWmlFGjKZIPD2aQcGOuLkuJcr4RGLSyyKqXai QKXKKUd8nG7xE9AH8VU4/4Y9melQRGJMruf6H/MVB6CjnCMUKVbKzohA5YYkAKJ3dIYr sDpRPEeyyAjGsL7NjfTUVzIlkOpMAriEnBi7Jd31UAKUAJp1DRpZj0gAUEEqEuHv+ATM 2IKOZyQs7+g2TF6MD62eplR+SakA4U8ev1BiCea+Un7SjoC/7LCq/OUz40sjoEVbvV42 0r3fBhcKURsbpBWlrghz5Ckr2CrPntBsW/vNneA12Y+H/8LGYl8ZQTgCrtqBD9hO0dMc KPUg== X-Gm-Message-State: ACgBeo2UkMHXSvrP/hvWUxqbn3FRllXfoaf/EaGtT7xuBxKb/ngAbdJc N3l3zG4reXyU7gtcYZ0oQJ+VdhKFtHeMvD/8JEnrLw== X-Google-Smtp-Source: AA6agR42WGqJyo3CQ/wsWtBH4QMjo2FwsKfgLb/Aa3jKv6m3xoE+xzdLy95JXttnzairFt+14Q8rzI0UFGuD4qoW950= X-Received: by 2002:a0d:d850:0:b0:340:d2c0:b022 with SMTP id a77-20020a0dd850000000b00340d2c0b022mr16165795ywe.469.1661959731749; Wed, 31 Aug 2022 08:28:51 -0700 (PDT) MIME-Version: 1.0 References: <20220830214919.53220-1-surenb@google.com> <20220831084230.3ti3vitrzhzsu3fs@moria.home.lan> <20220831101948.f3etturccmp5ovkl@suse.de> In-Reply-To: From: Suren Baghdasaryan Date: Wed, 31 Aug 2022 08:28:40 -0700 Message-ID: Subject: Re: [RFC PATCH 00/30] Code tagging framework and applications To: Michal Hocko Cc: Mel Gorman , Kent Overstreet , Peter Zijlstra , Andrew Morton , Vlastimil Babka , Johannes Weiner , Roman Gushchin , Davidlohr Bueso , Matthew Wilcox , "Liam R. Howlett" , David Vernet , Juri Lelli , Laurent Dufour , Peter Xu , David Hildenbrand , Jens Axboe , mcgrof@kernel.org, masahiroy@kernel.org, nathan@kernel.org, changbin.du@intel.com, ytcoode@gmail.com, Vincent Guittot , Dietmar Eggemann , Steven Rostedt , Benjamin Segall , Daniel Bristot de Oliveira , Valentin Schneider , Christopher Lameter , Pekka Enberg , Joonsoo Kim , 42.hyeyoo@gmail.com, Alexander Potapenko , Marco Elver , dvyukov@google.com, Shakeel Butt , Muchun Song , arnd@arndb.de, jbaron@akamai.com, David Rientjes , Minchan Kim , Kalesh Singh , kernel-team , linux-mm , iommu@lists.linux.dev, kasan-dev@googlegroups.com, io-uring@vger.kernel.org, linux-arch@vger.kernel.org, xen-devel@lists.xenproject.org, linux-bcache@vger.kernel.org, linux-modules@vger.kernel.org, LKML Content-Type: text/plain; charset="UTF-8" Precedence: bulk List-ID: X-Mailing-List: linux-bcache@vger.kernel.org On Wed, Aug 31, 2022 at 3:47 AM Michal Hocko wrote: > > On Wed 31-08-22 11:19:48, Mel Gorman wrote: > > On Wed, Aug 31, 2022 at 04:42:30AM -0400, Kent Overstreet wrote: > > > On Wed, Aug 31, 2022 at 09:38:27AM +0200, Peter Zijlstra wrote: > > > > On Tue, Aug 30, 2022 at 02:48:49PM -0700, Suren Baghdasaryan wrote: > > > > > =========================== > > > > > Code tagging framework > > > > > =========================== > > > > > Code tag is a structure identifying a specific location in the source code > > > > > which is generated at compile time and can be embedded in an application- > > > > > specific structure. Several applications of code tagging are included in > > > > > this RFC, such as memory allocation tracking, dynamic fault injection, > > > > > latency tracking and improved error code reporting. > > > > > Basically, it takes the old trick of "define a special elf section for > > > > > objects of a given type so that we can iterate over them at runtime" and > > > > > creates a proper library for it. > > > > > > > > I might be super dense this morning, but what!? I've skimmed through the > > > > set and I don't think I get it. > > > > > > > > What does this provide that ftrace/kprobes don't already allow? > > > > > > You're kidding, right? > > > > It's a valid question. From the description, it main addition that would > > be hard to do with ftrace or probes is catching where an error code is > > returned. A secondary addition would be catching all historical state and > > not just state since the tracing started. > > > > It's also unclear *who* would enable this. It looks like it would mostly > > have value during the development stage of an embedded platform to track > > kernel memory usage on a per-application basis in an environment where it > > may be difficult to setup tracing and tracking. Would it ever be enabled > > in production? Would a distribution ever enable this? If it's enabled, any > > overhead cannot be disabled/enabled at run or boot time so anyone enabling > > this would carry the cost without never necessarily consuming the data. Thank you for the question. For memory tracking my intent is to have a mechanism that can be enabled in the field testing (pre-production testing on a large population of internal users). The issue that we are often facing is when some memory leaks are happening in the field but very hard to reproduce locally. We get a bugreport from the user which indicates it but often has not enough information to track it. Note that quite often these leaks/issues happen in the drivers, so even simply finding out where they came from is a big help. The way I envision this mechanism to be used is to enable the basic memory tracking in the field tests and have a user space process collecting the allocation statistics periodically (say once an hour). Once it detects some counter growing infinitely or atypically (the definition of this is left to the user space) it can enable context capturing only for that specific location, still keeping the overhead to the minimum but getting more information about potential issues. Collected stats and contexts are then attached to the bugreport and we get more visibility into the issue when we receive it. The goal is to provide a mechanism with low enough overhead that it can be enabled all the time during these field tests without affecting the device's performance profiles. Tracing is very cheap when it's disabled but having it enabled all the time would introduce higher overhead than the counter manipulations. My apologies, I should have clarified all this in this cover letter from the beginning. As for other applications, maybe I'm not such an advanced user of tracing but I think only the latency tracking application might be done with tracing, assuming we have all the right tracepoints but I don't see how we would use tracing for fault injections and descriptive error codes. Again, I might be mistaken. Thanks, Suren. > > > > It might be an ease-of-use thing. Gathering the information from traces > > is tricky and would need combining multiple different elements and that > > is development effort but not impossible. > > > > Whatever asking for an explanation as to why equivalent functionality > > cannot not be created from ftrace/kprobe/eBPF/whatever is reasonable. > > Fully agreed and this is especially true for a change this size > 77 files changed, 3406 insertions(+), 703 deletions(-) > > -- > Michal Hocko > SUSE Labs