Merge branch 'perf-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull perf changes from Ingo Molnar: "Main changes: Kernel side changes: - Add SNB/IVB/HSW client uncore memory controller support (Stephane Eranian) - Fix various x86/P4 PMU driver bugs (Don Zickus) Tooling, user visible changes: - Add several futex 'perf bench' microbenchmarks (Davidlohr Bueso) - Speed up thread map generation (Don Zickus) - Introduce 'perf kvm --list-cmds' command line option for use by scripts (Ramkumar Ramachandra) - Print the evsel name in the annotate stdio output, prep to fix support outputting annotation for multiple events, not just for the first one (Arnaldo Carvalho de Melo) - Allow setting preferred callchain method in .perfconfig (Jiri Olsa) - Show in what binaries/modules 'perf probe's are set (Masami Hiramatsu) - Support distro-style debuginfo for uprobe in 'perf probe' (Masami Hiramatsu) Tooling, internal changes and fixes: - Use tid in mmap/mmap2 events to find maps (Don Zickus) - Record the reason for filtering an address_location (Namhyung Kim) - Apply all filters to an addr_location (Namhyung Kim) - Merge al->filtered with hist_entry->filtered in report/hists (Namhyung Kim) - Fix memory leak when synthesizing thread records (Namhyung Kim) - Use ui__has_annotation() in 'report' (Namhyung Kim) - hists browser refactorings to reuse code accross UIs (Namhyung Kim) - Add support for the new DWARF unwinder library in elfutils (Jiri Olsa) - Fix build race in the generation of bison files (Jiri Olsa) - Further streamline the feature detection display, trimming it a bit to show just the libraries detected, using VF=1 gets a more verbose output, showing the less interesting feature checks as well (Jiri Olsa). - Check compatible symtab type before loading dso (Namhyung Kim) - Check return value of filename__read_debuglink() (Stephane Eranian) - Move some hashing and fs related code from tools/perf/util/ to tools/lib/ so that it can be used by more tools/ living utilities (Borislav Petkov) - Prepare DWARF unwinding code for using an elfutils alternative unwinding library (Jiri Olsa) - Fix DWARF unwind max_stack processing (Jiri Olsa) - Add dwarf unwind 'perf test' entry (Jiri Olsa) - 'perf probe' improvements including memory leak fixes, sharing the intlist class with other tools, uprobes/kprobes code sharing and use of ref_reloc_sym (Masami Hiramatsu) - Shorten sample symbol resolving by adding cpumode to struct addr_location (Arnaldo Carvalho de Melo) - Fix synthesizing mmaps for threads (Don Zickus) - Fix invalid output on event group stdio report (Namhyung Kim) - Fixup header alignment in 'perf sched latency' output (Ramkumar Ramachandra) - Fix off-by-one error in 'perf timechart record' argv handling (Ramkumar Ramachandra) Tooling, cleanups: - Remove unused thread__find_map function (Jiri Olsa) - Remove unused simple_strtoul() function (Ramkumar Ramachandra) Tooling, documentation updates: - Update function names in debug messages (Ramkumar Ramachandra) - Update some code references in design.txt (Ramkumar Ramachandra) - Clarify load-latency information in the 'perf mem' docs (Andi Kleen) - Clarify x86 register naming in 'perf probe' docs (Andi Kleen)" * 'perf-core-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (96 commits) perf tools: Remove unused simple_strtoul() function perf tools: Update some code references in design.txt perf evsel: Update function names in debug messages perf tools: Remove thread__find_map function perf annotate: Print the evsel name in the stdio output perf report: Use ui__has_annotation() perf tools: Fix memory leak when synthesizing thread records perf tools: Use tid in mmap/mmap2 events to find maps perf report: Merge al->filtered with hist_entry->filtered perf symbols: Apply all filters to an addr_location perf symbols: Record the reason for filtering an address_location perf sched: Fixup header alignment in 'latency' output perf timechart: Fix off-by-one error in 'record' argv handling perf machine: Factor machine__find_thread to take tid argument perf tools: Speed up thread map generation perf kvm: introduce --list-cmds for use by scripts perf ui hists: Pass evsel to hpp->header/width functions explicitly perf symbols: Introduce thread__find_cpumode_addr_location perf session: Change header.misc dump from decimal to hex perf ui/tui: Reuse generic __hpp__fmt() code ...
author: Linus Torvalds <torvalds@linux-foundation.org> 2014-03-31 14:13:25 -0400
committer: Linus Torvalds <torvalds@linux-foundation.org> 2014-03-31 14:13:25 -0400
commit: 8c292f11744297dfb3a69f4a0bccbe4a6417b50d (patch)
tree: f1a89560de25a69b697d459a9b5cf2e738038d9f /kernel
parent: d31605dc8a63f1df28443ddb3560b1079417af92 (diff)
parent: 538592ff0b008237ae88f5ce5fb1247127dc3ce5 (diff)
3 files changed, 51 insertions, 16 deletions
diff --git a/kernel/events/core.c b/kernel/events/core.c
index fa0b2d4ad83c..661951ab8ae7 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -231,11 +231,29 @@ int perf_cpu_time_max_percent_handler(struct ctl_table *table, int write,
 #define NR_ACCUMULATED_SAMPLES 128
 static DEFINE_PER_CPU(u64, running_sample_length);
-void perf_sample_event_took(u64 sample_len_ns)
+static void perf_duration_warn(struct irq_work *w)
 {
+        u64 allowed_ns = ACCESS_ONCE(perf_sample_allowed_ns);
        u64 avg_local_sample_len;
        u64 local_samples_len;
+        local_samples_len = __get_cpu_var(running_sample_length);
+        avg_local_sample_len = local_samples_len/NR_ACCUMULATED_SAMPLES;
+        printk_ratelimited(KERN_WARNING
+                        "perf interrupt took too long (%lld > %lld), lowering "
+                        "kernel.perf_event_max_sample_rate to %d\n",
+                        avg_local_sample_len, allowed_ns >> 1,
+                        sysctl_perf_event_sample_rate);
+}
+static DEFINE_IRQ_WORK(perf_duration_work, perf_duration_warn);
+void perf_sample_event_took(u64 sample_len_ns)
+{
        u64 allowed_ns = ACCESS_ONCE(perf_sample_allowed_ns);
+        u64 avg_local_sample_len;
+        u64 local_samples_len;
        if (allowed_ns == 0)
                return;
@@ -263,13 +281,14 @@ void perf_sample_event_took(u64 sample_len_ns)
        sysctl_perf_event_sample_rate = max_samples_per_tick * HZ;
        perf_sample_period_ns = NSEC_PER_SEC / sysctl_perf_event_sample_rate;
-        printk_ratelimited(KERN_WARNING
-                        "perf samples too long (%lld > %lld), lowering "
-                        "kernel.perf_event_max_sample_rate to %d\n",
-                        avg_local_sample_len, allowed_ns,
-                        sysctl_perf_event_sample_rate);
        update_perf_cpu_limits();
+        if (!irq_work_queue(&perf_duration_work)) {
+                early_printk("perf interrupt took too long (%lld > %lld), lowering "
+                             "kernel.perf_event_max_sample_rate to %d\n",
+                             avg_local_sample_len, allowed_ns >> 1,
+                             sysctl_perf_event_sample_rate);
+        }
 }
 static atomic64_t perf_event_id;
@@ -1714,7 +1733,7 @@ group_sched_in(struct perf_event *group_event,
               struct perf_event_context *ctx)
 {
        struct perf_event *event, *partial_group = NULL;
-        struct pmu *pmu = group_event->pmu;
+        struct pmu *pmu = ctx->pmu;
        u64 now = ctx->time;
        bool simulate = false;
@@ -2563,8 +2582,6 @@ static void perf_branch_stack_sched_in(struct task_struct *prev,
                if (cpuctx->ctx.nr_branch_stack > 0
                    && pmu->flush_branch_stack) {
-                        pmu = cpuctx->ctx.pmu;
                        perf_ctx_lock(cpuctx, cpuctx->task_ctx);
                        perf_pmu_disable(pmu);
@@ -6294,7 +6311,7 @@ static int perf_event_idx_default(struct perf_event *event)
 * Ensures all contexts with the same task_ctx_nr have the same
 * pmu_cpu_context too.
 */
-static void *find_pmu_context(int ctxn)
+static struct perf_cpu_context __percpu *find_pmu_context(int ctxn)
 {
        struct pmu *pmu;
diff --git a/kernel/irq_work.c b/kernel/irq_work.c
index 55fcce6065cf..a82170e2fa78 100644
--- a/kernel/irq_work.c
+++ b/kernel/irq_work.c
@@ -61,11 +61,11 @@ void __weak arch_irq_work_raise(void)
 *
 * Can be re-enqueued while the callback is still in progress.
 */
-void irq_work_queue(struct irq_work *work)
+bool irq_work_queue(struct irq_work *work)
 {
        /* Only queue if not already pending */
        if (!irq_work_claim(work))
-                return;
+                return false;
        /* Queue the entry and raise the IPI if needed. */
        preempt_disable();
@@ -83,6 +83,8 @@ void irq_work_queue(struct irq_work *work)
        }
        preempt_enable();
+        return true;
 }
 EXPORT_SYMBOL_GPL(irq_work_queue);
diff --git a/kernel/trace/trace_event_perf.c b/kernel/trace/trace_event_perf.c
index e854f420e033..c894614de14d 100644
--- a/kernel/trace/trace_event_perf.c
+++ b/kernel/trace/trace_event_perf.c
@@ -31,9 +31,25 @@ static int perf_trace_event_perm(struct ftrace_event_call *tp_event,
        }
        /* The ftrace function trace is allowed only for root. */
-        if (ftrace_event_is_function(tp_event) &&
+        if (ftrace_event_is_function(tp_event)) {
-            perf_paranoid_tracepoint_raw() && !capable(CAP_SYS_ADMIN))
+                if (perf_paranoid_tracepoint_raw() && !capable(CAP_SYS_ADMIN))
-                return -EPERM;
+                        return -EPERM;
+                /*
+                 * We don't allow user space callchains for  function trace
+                 * event, due to issues with page faults while tracing page
+                 * fault handler and its overall trickiness nature.
+                 */
+                if (!p_event->attr.exclude_callchain_user)
+                        return -EINVAL;
+                /*
+                 * Same reason to disable user stack dump as for user space
+                 * callchains above.
+                 */
+                if (p_event->attr.sample_type & PERF_SAMPLE_STACK_USER)
+                        return -EINVAL;
+        }
        /* No tracing, just counting, so no obvious leak */
        if (!(p_event->attr.sample_type & PERF_SAMPLE_RAW))
author	Linus Torvalds <torvalds@linux-foundation.org>	2014-03-31 14:13:25 -0400
committer	Linus Torvalds <torvalds@linux-foundation.org>	2014-03-31 14:13:25 -0400
commit	8c292f11744297dfb3a69f4a0bccbe4a6417b50d (patch)
tree	f1a89560de25a69b697d459a9b5cf2e738038d9f /kernel
parent	d31605dc8a63f1df28443ddb3560b1079417af92 (diff)
parent	538592ff0b008237ae88f5ce5fb1247127dc3ce5 (diff)