Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1591705 > unrolled thread

[PATCH v3 0/5] coresight: enable debug module

Started byLeo Yan <leo.yan@linaro.org>
First post2017-03-03 08:10 +0100
Last post2017-03-09 14:30 +0100
Articles 14 — 5 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v3 0/5] coresight: enable debug module Leo Yan <leo.yan@linaro.org> - 2017-03-03 08:10 +0100
    [PATCH v3 2/5] coresight: refactor with function of_coresight_get_cpu Leo Yan <leo.yan@linaro.org> - 2017-03-03 08:10 +0100
    [PATCH v3 5/5] arm64: dts: hi6220: register debug module Leo Yan <leo.yan@linaro.org> - 2017-03-03 09:00 +0100
    [PATCH v3 4/5] clk: hi6220: add debug APB clock Leo Yan <leo.yan@linaro.org> - 2017-03-03 10:10 +0100
      Re: [PATCH v3 4/5] clk: hi6220: add debug APB clock Stephen Boyd <sboyd@codeaurora.org> - 2017-03-04 01:10 +0100
    [PATCH v3 3/5] coresight: add support for debug module Leo Yan <leo.yan@linaro.org> - 2017-03-03 13:40 +0100
      Re: [v3 3/5] coresight: add support for debug module Suzuki K Poulose <Suzuki.Poulose@arm.com> - 2017-03-09 18:10 +0100
        Re: [v3 3/5] coresight: add support for debug module Leo Yan <leo.yan@linaro.org> - 2017-03-09 19:10 +0100
          Re: [v3 3/5] coresight: add support for debug module Suzuki K Poulose <Suzuki.Poulose@arm.com> - 2017-03-10 15:40 +0100
            Re: [v3 3/5] coresight: add support for debug module Leo Yan <leo.yan@linaro.org> - 2017-03-13 09:20 +0100
            Re: [v3 3/5] coresight: add support for debug module Mathieu Poirier <mathieu.poirier@linaro.org> - 2017-03-13 18:00 +0100
          Re: [v3 3/5] coresight: add support for debug module Mathieu Poirier <mathieu.poirier@linaro.org> - 2017-03-13 17:40 +0100
    [PATCH v3 1/5] coresight: bindings for debug module Leo Yan <leo.yan@linaro.org> - 2017-03-03 19:50 +0100
      Re: [v3 1/5] coresight: bindings for debug module Suzuki K Poulose <suzuki.poulose@arm.com> - 2017-03-09 14:30 +0100

#1591705 — [PATCH v3 0/5] coresight: enable debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 08:10 +0100
Subject[PATCH v3 0/5] coresight: enable debug module
Message-ID<tgPvH-5Zw-11@gated-at.bofh.it>
ARMv8 architecture reference manual (ARM DDI 0487A.k) Chapter H7 "The
Sample-based Profiling Extension" has description for sampling
registers, we can utilize these registers to check program counter
value with combined CPU exception level, secure state, etc. So this is
helpful for CPU lockup bugs, e.g. if one CPU has run into infinite loop
with IRQ disabled; the 'hang' CPU cannot switch context and handle any
interrupt, so it cannot handle SMP call for stack dump, etc.

This patch series is to enable coresight debug module with sample-based
registers and register call back notifier for PCSR register dumping
when panic happens, so we can see below dumping info for panic; and
this patch series has considered the conditions for access permission
for debug registers self, so this can avoid access debug registers when
CPU power domain is off; the driver also try to figure out the CPU is
in secure or non-secure state.

The last two patches in this series is to enable debug unit on 96boards
Hikey, the first patch is to add apb clock for debug unit and the second
patch is to add DT nodes for debug unit. As result we can below log
after input command: echo c > /proc/sysrq-trigger:

ARM external debug module:
CPU[0]:
 EDPRSR:  0000000b (Power:On DLK:Unlock)
 EDPCSR:  [<ffff00000808eb54>] handle_IPI+0xe4/0x150
 EDCIDSR: 00000000
 EDVIDSR: 90000000 (State:Non-secure Mode:EL1/0 Width:64bits VMID:0)
CPU[1]:
 EDPRSR:  0000000b (Power:On DLK:Unlock)
 EDPCSR:  [<ffff0000087a64c0>] debug_notifier_call+0x108/0x288
 EDCIDSR: 00000000
 EDVIDSR: 90000000 (State:Non-secure Mode:EL1/0 Width:64bits VMID:0)

[...]

Changes from v2:
* According to Mathieu Poirier suggestion, applied some minor fixes.
* Added two extra patches for enabling debug module on Hikey.

Changes from v1:
* According to Mike Leach suggestion, removed the binding for debug
  module clocks which have been directly provided by CPU clocks.
* According to Mathieu Poirier suggestion, added function
  of_coresight_get_cpu() and some minor refactors for debug module
  driver.

Changes from RFC:
* According to Mike Leach suggestion, added check for EDPRSR to avoid
  lockup; added supporting EDVIDSR and EDCIDSR registers.
* According to Mark Rutland and Mathieu Poirier suggestion, rewrote
  the documentation for DT binding.
* According to Mark and Mathieu suggestion, refined debug driver.


Leo Yan (5):
  coresight: bindings for debug module
  coresight: refactor with function of_coresight_get_cpu
  coresight: add support for debug module
  clk: hi6220: add debug APB clock
  arm64: dts: hi6220: register debug module

 .../devicetree/bindings/arm/coresight-debug.txt    |  40 +++
 arch/arm64/boot/dts/hisilicon/hi6220.dtsi          |  64 ++++
 drivers/clk/hisilicon/clk-hi6220.c                 |   1 +
 drivers/hwtracing/coresight/Kconfig                |  10 +
 drivers/hwtracing/coresight/Makefile               |   1 +
 drivers/hwtracing/coresight/coresight-debug.c      | 377 +++++++++++++++++++++
 drivers/hwtracing/coresight/of_coresight.c         |  37 +-
 include/dt-bindings/clock/hi6220-clock.h           |   5 +-
 include/linux/coresight.h                          |   2 +
 9 files changed, 524 insertions(+), 13 deletions(-)
 create mode 100644 Documentation/devicetree/bindings/arm/coresight-debug.txt
 create mode 100644 drivers/hwtracing/coresight/coresight-debug.c

-- 
2.7.4

[toc] | [next] | [standalone]


#1591706 — [PATCH v3 2/5] coresight: refactor with function of_coresight_get_cpu

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 08:10 +0100
Subject[PATCH v3 2/5] coresight: refactor with function of_coresight_get_cpu
Message-ID<tgPvH-5Zw-9@gated-at.bofh.it>
In reply to#1591705
This is refactor to add function of_coresight_get_cpu(), so it's used to
retrieve CPU id for coresight component. Finally can use it as a common
function for multiple places.

Suggested-by: Mathieu Poirier <mathieu.poirier@linaro.org>
Signed-off-by: Leo Yan <leo.yan@linaro.org>
---
 drivers/hwtracing/coresight/of_coresight.c | 37 ++++++++++++++++++++----------
 include/linux/coresight.h                  |  2 ++
 2 files changed, 27 insertions(+), 12 deletions(-)

diff --git a/drivers/hwtracing/coresight/of_coresight.c b/drivers/hwtracing/coresight/of_coresight.c
index 629e031..d9a12fb 100644
--- a/drivers/hwtracing/coresight/of_coresight.c
+++ b/drivers/hwtracing/coresight/of_coresight.c
@@ -101,14 +101,36 @@ static int of_coresight_alloc_memory(struct device *dev,
 	return 0;
 }
 
+int of_coresight_get_cpu(struct device_node *node)
+{
+	int cpu;
+	struct device_node *dn;
+
+	dn = of_parse_phandle(node, "cpu", 0);
+
+	/* Affinity defaults to CPU0 */
+	if (!dn)
+		return 0;
+
+	for_each_possible_cpu(cpu) {
+		if (dn == of_get_cpu_node(cpu, NULL)) {
+			of_node_put(dn);
+			return cpu;
+		}
+	}
+
+	/* Affinity to CPU0 if no cpu nodes are found */
+	of_node_put(dn);
+	return 0;
+}
+
 struct coresight_platform_data *of_get_coresight_platform_data(
 				struct device *dev, struct device_node *node)
 {
-	int i = 0, ret = 0, cpu;
+	int i = 0, ret = 0;
 	struct coresight_platform_data *pdata;
 	struct of_endpoint endpoint, rendpoint;
 	struct device *rdev;
-	struct device_node *dn;
 	struct device_node *ep = NULL;
 	struct device_node *rparent = NULL;
 	struct device_node *rport = NULL;
@@ -175,16 +197,7 @@ struct coresight_platform_data *of_get_coresight_platform_data(
 		} while (ep);
 	}
 
-	/* Affinity defaults to CPU0 */
-	pdata->cpu = 0;
-	dn = of_parse_phandle(node, "cpu", 0);
-	for (cpu = 0; dn && cpu < nr_cpu_ids; cpu++) {
-		if (dn == of_get_cpu_node(cpu, NULL)) {
-			pdata->cpu = cpu;
-			break;
-		}
-	}
-	of_node_put(dn);
+	pdata->cpu = of_coresight_get_cpu(node);
 
 	return pdata;
 }
diff --git a/include/linux/coresight.h b/include/linux/coresight.h
index 2a5982c..7b29743 100644
--- a/include/linux/coresight.h
+++ b/include/linux/coresight.h
@@ -263,9 +263,11 @@ static inline int coresight_timeout(void __iomem *addr, u32 offset,
 #endif
 
 #ifdef CONFIG_OF
+extern int of_coresight_get_cpu(struct device_node *node);
 extern struct coresight_platform_data *of_get_coresight_platform_data(
 				struct device *dev, struct device_node *node);
 #else
+static int of_coresight_get_cpu(struct device_node *node) { return 0; }
 static inline struct coresight_platform_data *of_get_coresight_platform_data(
 	struct device *dev, struct device_node *node) { return NULL; }
 #endif
-- 
2.7.4

[toc] | [prev] | [next] | [standalone]


#1591727 — [PATCH v3 5/5] arm64: dts: hi6220: register debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 09:00 +0100
Subject[PATCH v3 5/5] arm64: dts: hi6220: register debug module
Message-ID<tgQi5-6mq-1@gated-at.bofh.it>
In reply to#1591705
Bind debug module driver for Hi6220.

Signed-off-by: Leo Yan <leo.yan@linaro.org>
---
 arch/arm64/boot/dts/hisilicon/hi6220.dtsi | 64 +++++++++++++++++++++++++++++++
 1 file changed, 64 insertions(+)

diff --git a/arch/arm64/boot/dts/hisilicon/hi6220.dtsi b/arch/arm64/boot/dts/hisilicon/hi6220.dtsi
index 470461d..ed271ed 100644
--- a/arch/arm64/boot/dts/hisilicon/hi6220.dtsi
+++ b/arch/arm64/boot/dts/hisilicon/hi6220.dtsi
@@ -913,5 +913,69 @@
 				};
 			};
 		};
+
+		debug@f6590000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf6590000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu0>;
+		};
+
+		debug@f6592000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf6592000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu1>;
+		};
+
+		debug@f6594000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf6594000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu2>;
+		};
+
+		debug@f6596000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf6596000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu3>;
+		};
+
+		debug@f65d0000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf65d0000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu4>;
+		};
+
+		debug@f65d2000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf65d2000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu5>;
+		};
+
+		debug@f65d4000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf65d4000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu6>;
+		};
+
+		debug@f65d6000 {
+			compatible = "arm,coresight-debug","arm,primecell";
+			reg = <0 0xf65d6000 0 0x1000>;
+			clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+			clock-names = "apb_pclk";
+			cpu = <&cpu7>;
+		};
 	};
 };
-- 
2.7.4

[toc] | [prev] | [next] | [standalone]


#1591775 — [PATCH v3 4/5] clk: hi6220: add debug APB clock

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 10:10 +0100
Subject[PATCH v3 4/5] clk: hi6220: add debug APB clock
Message-ID<tgRnR-7or-41@gated-at.bofh.it>
In reply to#1591705
The debug APB clock is absent in hi6220 driver, so this patch is to add
support for it.

Signed-off-by: Leo Yan <leo.yan@linaro.org>
---
 drivers/clk/hisilicon/clk-hi6220.c       | 1 +
 include/dt-bindings/clock/hi6220-clock.h | 5 ++++-
 2 files changed, 5 insertions(+), 1 deletion(-)

diff --git a/drivers/clk/hisilicon/clk-hi6220.c b/drivers/clk/hisilicon/clk-hi6220.c
index c0e8e1f..6879c1f 100644
--- a/drivers/clk/hisilicon/clk-hi6220.c
+++ b/drivers/clk/hisilicon/clk-hi6220.c
@@ -134,6 +134,7 @@ static struct hisi_gate_clock hi6220_separated_gate_clks_sys[] __initdata = {
 	{ HI6220_UART4_PCLK,    "uart4_pclk",    "uart4_src",      CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x230, 8,  0, },
 	{ HI6220_SPI_CLK,       "spi_clk",       "clk_150m",       CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x230, 9,  0, },
 	{ HI6220_TSENSOR_CLK,   "tsensor_clk",   "clk_bus",        CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x230, 12, 0, },
+	{ HI6220_DAPB_CLK,      "dapb_clk",      "cs_dapb",        CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x230, 18, 0, },
 	{ HI6220_MMU_CLK,       "mmu_clk",       "ddrc_axi1",      CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x240, 11, 0, },
 	{ HI6220_HIFI_SEL,      "hifi_sel",      "hifi_src",       CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x270, 0,  0, },
 	{ HI6220_MMC0_SYSPLL,   "mmc0_syspll",   "syspll",         CLK_SET_RATE_PARENT|CLK_IGNORE_UNUSED, 0x270, 1,  0, },
diff --git a/include/dt-bindings/clock/hi6220-clock.h b/include/dt-bindings/clock/hi6220-clock.h
index 6b03c84..b8ba665 100644
--- a/include/dt-bindings/clock/hi6220-clock.h
+++ b/include/dt-bindings/clock/hi6220-clock.h
@@ -124,7 +124,10 @@
 #define HI6220_CS_DAPB		57
 #define HI6220_CS_ATB_DIV	58
 
-#define HI6220_SYS_NR_CLKS	59
+/* gate clock */
+#define HI6220_DAPB_CLK		59
+
+#define HI6220_SYS_NR_CLKS	60
 
 /* clk in Hi6220 media controller */
 /* gate clocks */
-- 
2.7.4

[toc] | [prev] | [next] | [standalone]


#1592349 — Re: [PATCH v3 4/5] clk: hi6220: add debug APB clock

FromStephen Boyd <sboyd@codeaurora.org>
Date2017-03-04 01:10 +0100
SubjectRe: [PATCH v3 4/5] clk: hi6220: add debug APB clock
Message-ID<th5qN-nH-3@gated-at.bofh.it>
In reply to#1591775
On 03/03, Leo Yan wrote:
> The debug APB clock is absent in hi6220 driver, so this patch is to add
> support for it.
> 
> Signed-off-by: Leo Yan <leo.yan@linaro.org>
> ---

Acked-by: Stephen Boyd <sboyd@codeaurora.org>

-- 
Qualcomm Innovation Center, Inc. is a member of Code Aurora Forum,
a Linux Foundation Collaborative Project

[toc] | [prev] | [next] | [standalone]


#1591908 — [PATCH v3 3/5] coresight: add support for debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 13:40 +0100
Subject[PATCH v3 3/5] coresight: add support for debug module
Message-ID<tgUF3-1cK-1@gated-at.bofh.it>
In reply to#1591705
Coresight includes debug module and usually the module connects with CPU
debug logic. ARMv8 architecture reference manual (ARM DDI 0487A.k) has
description for related info in "Part H: External Debug".

Chapter H7 "The Sample-based Profiling Extension" introduces several
sampling registers, e.g. we can check program counter value with
combined CPU exception level, secure state, etc. So this is helpful for
analysis CPU lockup scenarios, e.g. if one CPU has run into infinite
loop with IRQ disabled. In this case the CPU cannot switch context and
handle any interrupt (including IPIs), as the result it cannot handle
SMP call for stack dump.

This patch is to enable coresight debug module, so firstly this driver
is to bind apb clock for debug module and this is to ensure the debug
module can be accessed from program or external debugger. And the driver
uses sample-based registers for debug purpose, e.g. when system detects
the CPU lockup and trigger panic, the driver will dump program counter
and combined context registers (EDCIDSR, EDVIDSR); by parsing context
registers so can quickly get to know CPU secure state, exception level,
etc.

Some of the debug module registers are located in CPU power domain, so
in the driver it has checked the power state for CPU before accessing
registers within CPU power domain. For most safe way to use this driver,
it's suggested to disable CPU low power states, this can simply set
"nohlt" in kernel command line.

Signed-off-by: Leo Yan <leo.yan@linaro.org>
---
 drivers/hwtracing/coresight/Kconfig           |  10 +
 drivers/hwtracing/coresight/Makefile          |   1 +
 drivers/hwtracing/coresight/coresight-debug.c | 377 ++++++++++++++++++++++++++
 3 files changed, 388 insertions(+)
 create mode 100644 drivers/hwtracing/coresight/coresight-debug.c

diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
index 130cb21..3ed651e 100644
--- a/drivers/hwtracing/coresight/Kconfig
+++ b/drivers/hwtracing/coresight/Kconfig
@@ -89,4 +89,14 @@ config CORESIGHT_STM
 	  logging useful software events or data coming from various entities
 	  in the system, possibly running different OSs
 
+config CORESIGHT_DEBUG
+	bool "CoreSight debug driver"
+	depends on ARM || ARM64
+	help
+	  This driver provides support for coresight debugging module. This
+	  is primarily used to dump sample-based profiling registers for
+	  panic. To avoid lockups when accessing debug module registers,
+	  it is safer to disable CPU low power states (like "nohlt" on the
+	  kernel command line) when using this feature.
+
 endif
diff --git a/drivers/hwtracing/coresight/Makefile b/drivers/hwtracing/coresight/Makefile
index af480d9..d540d45 100644
--- a/drivers/hwtracing/coresight/Makefile
+++ b/drivers/hwtracing/coresight/Makefile
@@ -16,3 +16,4 @@ obj-$(CONFIG_CORESIGHT_SOURCE_ETM4X) += coresight-etm4x.o \
 					coresight-etm4x-sysfs.o
 obj-$(CONFIG_CORESIGHT_QCOM_REPLICATOR) += coresight-replicator-qcom.o
 obj-$(CONFIG_CORESIGHT_STM) += coresight-stm.o
+obj-$(CONFIG_CORESIGHT_DEBUG) += coresight-debug.o
diff --git a/drivers/hwtracing/coresight/coresight-debug.c b/drivers/hwtracing/coresight/coresight-debug.c
new file mode 100644
index 0000000..9553fb2
--- /dev/null
+++ b/drivers/hwtracing/coresight/coresight-debug.c
@@ -0,0 +1,377 @@
+/*
+ * Copyright (c) 2017 Linaro Limited. All rights reserved.
+ *
+ * Author: Leo Yan <leo.yan@linaro.org>
+ *
+ * This program is free software; you can redistribute it and/or modify it
+ * under the terms of the GNU General Public License version 2 as published by
+ * the Free Software Foundation.
+ *
+ * This program is distributed in the hope that it will be useful, but WITHOUT
+ * ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or
+ * FITNESS FOR A PARTICULAR PURPOSE.  See the GNU General Public License for
+ * more details.
+ *
+ * You should have received a copy of the GNU General Public License along with
+ * this program.  If not, see <http://www.gnu.org/licenses/>.
+ *
+ */
+#include <linux/amba/bus.h>
+#include <linux/coresight.h>
+#include <linux/cpu.h>
+#include <linux/device.h>
+#include <linux/err.h>
+#include <linux/init.h>
+#include <linux/io.h>
+#include <linux/kernel.h>
+#include <linux/moduleparam.h>
+#include <linux/slab.h>
+#include <linux/smp.h>
+#include <linux/types.h>
+#include <linux/uaccess.h>
+
+#include "coresight-priv.h"
+
+#define EDPCSR				0x0A0
+#define EDCIDSR				0x0A4
+#define EDVIDSR				0x0A8
+#define EDPCSR_HI			0x0AC
+#define EDOSLAR				0x300
+#define EDPRSR				0x314
+#define EDDEVID1			0xFC4
+#define EDDEVID				0xFC8
+
+#define EDPCSR_PROHIBITED		0xFFFFFFFF
+
+/* bits definition for EDPCSR */
+#ifndef CONFIG_64BIT
+#define EDPCSR_THUMB			BIT(0)
+#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
+#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)
+#endif
+
+/* bits definition for EDPRSR */
+#define EDPRSR_DLK			BIT(6)
+#define EDPRSR_PU			BIT(0)
+
+/* bits definition for EDVIDSR */
+#define EDVIDSR_NS			BIT(31)
+#define EDVIDSR_E2			BIT(30)
+#define EDVIDSR_E3			BIT(29)
+#define EDVIDSR_HV			BIT(28)
+#define EDVIDSR_VMID			GENMASK(7, 0)
+
+/* bits definition for EDDEVID1 */
+#define EDDEVID1_PCSR_OFFSET_MASK	GENMASK(3, 0)
+#define EDDEVID1_PCSR_OFFSET_INS_SET	(0x0)
+
+/* bits definition for EDDEVID */
+#define EDDEVID_PCSAMPLE_MODE		GENMASK(3, 0)
+#define EDDEVID_IMPL_EDPCSR_EDCIDSR	(0x2)
+#define EDDEVID_IMPL_FULL		(0x3)
+
+struct debug_drvdata {
+	void __iomem	*base;
+	struct device	*dev;
+	int		cpu;
+
+	bool		edpcsr_present;
+	bool		edvidsr_present;
+	bool		pc_has_offset;
+
+	u32		eddevid;
+	u32		eddevid1;
+
+	u32		edpcsr;
+	u32		edpcsr_hi;
+	u32		edprsr;
+	u32		edvidsr;
+	u32		edcidsr;
+};
+
+static DEFINE_PER_CPU(struct debug_drvdata *, debug_drvdata);
+
+static void debug_os_unlock(struct debug_drvdata *drvdata)
+{
+	/* Unlocks the debug registers */
+	writel_relaxed(0x0, drvdata->base + EDOSLAR);
+	wmb();
+}
+
+/*
+ * According to ARM DDI 0487A.k, before access external debug
+ * registers should firstly check the access permission; if any
+ * below condition has been met then cannot access debug
+ * registers to avoid lockup issue:
+ *
+ * - CPU power domain is powered off;
+ * - The OS Double Lock is locked;
+ *
+ * By checking EDPRSR can get to know if meet these conditions.
+ */
+static bool debug_access_permitted(struct debug_drvdata *drvdata)
+{
+	/* CPU is powered off */
+	if (!(drvdata->edprsr & EDPRSR_PU))
+		return false;
+
+	/* The OS Double Lock is locked */
+	if (drvdata->edprsr & EDPRSR_DLK)
+		return false;
+
+	return true;
+}
+
+static void debug_read_regs(struct debug_drvdata *drvdata)
+{
+	drvdata->edprsr = readl_relaxed(drvdata->base + EDPRSR);
+
+	if (!debug_access_permitted(drvdata))
+		return;
+
+	if (!drvdata->edpcsr_present)
+		return;
+
+	CS_UNLOCK(drvdata->base);
+
+	debug_os_unlock(drvdata);
+
+	drvdata->edpcsr = readl_relaxed(drvdata->base + EDPCSR);
+
+	/*
+	 * As described in ARM DDI 0487A.k, if the processing
+	 * element (PE) is in debug state, or sample-based
+	 * profiling is prohibited, EDPCSR reads as 0xFFFFFFFF;
+	 * EDCIDSR, EDVIDSR and EDPCSR_HI registers also become
+	 * UNKNOWN state. So directly bail out for this case.
+	 */
+	if (drvdata->edpcsr == EDPCSR_PROHIBITED) {
+		CS_LOCK(drvdata->base);
+		return;
+	}
+
+	/*
+	 * A read of the EDPCSR normally has the side-effect of
+	 * indirectly writing to EDCIDSR, EDVIDSR and EDPCSR_HI;
+	 * at this point it's safe to read value from them.
+	 */
+	drvdata->edcidsr = readl_relaxed(drvdata->base + EDCIDSR);
+#ifdef CONFIG_64BIT
+	drvdata->edpcsr_hi = readl_relaxed(drvdata->base + EDPCSR_HI);
+#endif
+
+	if (drvdata->edvidsr_present)
+		drvdata->edvidsr = readl_relaxed(drvdata->base + EDVIDSR);
+
+	CS_LOCK(drvdata->base);
+}
+
+#ifndef CONFIG_64BIT
+static bool debug_pc_has_offset(struct debug_drvdata *drvdata)
+{
+	u32 pcsr_offset;
+
+	pcsr_offset = drvdata->eddevid1 & EDDEVID1_PCSR_OFFSET_MASK;
+
+	return (pcsr_offset == EDDEVID1_PCSR_OFFSET_INS_SET);
+}
+
+static unsigned long debug_adjust_pc(struct debug_drvdata *drvdata,
+				     unsigned long pc)
+{
+	unsigned long arm_inst_offset = 0, thumb_inst_offset = 0;
+
+	if (debug_pc_has_offset(drvdata)) {
+		arm_inst_offset = 8;
+		thumb_inst_offset = 4;
+	}
+
+	/* Handle thumb instruction */
+	if (pc & EDPCSR_THUMB) {
+		pc = (pc & EDPCSR_THUMB_INST_MASK) - thumb_inst_offset;
+		return pc;
+	}
+
+	/*
+	 * Handle arm instruction offset, if the arm instruction
+	 * is not 4 byte alignment then it's possible the case
+	 * for implementation defined; keep original value for this
+	 * case and print info for notice.
+	 */
+	if (pc & BIT(1))
+		pr_emerg("Instruction offset is implementation defined\n");
+	else
+		pc = (pc & EDPCSR_ARM_INST_MASK) - arm_inst_offset;
+
+	return pc;
+}
+#endif
+
+static void debug_dump_regs(struct debug_drvdata *drvdata)
+{
+	unsigned long pc;
+
+	pr_emerg("\tEDPRSR:  %08x (Power:%s DLK:%s)\n", drvdata->edprsr,
+		 drvdata->edprsr & EDPRSR_PU ? "On" : "Off",
+		 drvdata->edprsr & EDPRSR_DLK ? "Lock" : "Unlock");
+
+	if (!debug_access_permitted(drvdata) || !drvdata->edpcsr_present) {
+		pr_emerg("No permission to access debug registers!\n");
+		return;
+	}
+
+	if (drvdata->edpcsr == EDPCSR_PROHIBITED) {
+		pr_emerg("CPU is in Debug state or profiling is prohibited!\n");
+		return;
+	}
+
+#ifdef CONFIG_64BIT
+	pc = (unsigned long)drvdata->edpcsr_hi << 32 |
+	     (unsigned long)drvdata->edpcsr;
+#else
+	pc = debug_adjust_pc(drvdata, (unsigned long)drvdata->edpcsr);
+#endif
+
+	pr_emerg("\tEDPCSR:  [<%p>] %pS\n", (void *)pc, (void *)pc);
+	pr_emerg("\tEDCIDSR: %08x\n", drvdata->edcidsr);
+
+	if (!drvdata->edvidsr_present)
+		return;
+
+	pr_emerg("\tEDVIDSR: %08x (State:%s Mode:%s Width:%s VMID:%x)\n",
+		 drvdata->edvidsr,
+		 drvdata->edvidsr & EDVIDSR_NS ? "Non-secure" : "Secure",
+		 drvdata->edvidsr & EDVIDSR_E3 ? "EL3" :
+			(drvdata->edvidsr & EDVIDSR_E2 ? "EL2" : "EL1/0"),
+		 drvdata->edvidsr & EDVIDSR_HV ? "64bits" : "32bits",
+		 drvdata->edvidsr & (u32)EDVIDSR_VMID);
+}
+
+/*
+ * Dump out information on panic.
+ */
+static int debug_notifier_call(struct notifier_block *self,
+			       unsigned long v, void *p)
+{
+	int cpu;
+
+	pr_emerg("ARM external debug module:\n");
+
+	for_each_possible_cpu(cpu) {
+		if (!per_cpu(debug_drvdata, cpu))
+			continue;
+
+		pr_emerg("CPU[%d]:\n", per_cpu(debug_drvdata, cpu)->cpu);
+
+		debug_read_regs(per_cpu(debug_drvdata, cpu));
+		debug_dump_regs(per_cpu(debug_drvdata, cpu));
+	}
+
+	return 0;
+}
+
+static struct notifier_block debug_notifier = {
+	.notifier_call = debug_notifier_call,
+};
+
+static void debug_init_arch_data(void *info)
+{
+	struct debug_drvdata *drvdata = info;
+	u32 mode;
+
+	CS_UNLOCK(drvdata->base);
+
+	debug_os_unlock(drvdata);
+
+	/* Read device info */
+	drvdata->eddevid  = readl_relaxed(drvdata->base + EDDEVID);
+	drvdata->eddevid1 = readl_relaxed(drvdata->base + EDDEVID1);
+
+	/* Parse implementation feature */
+	mode = drvdata->eddevid & EDDEVID_PCSAMPLE_MODE;
+	if (mode == EDDEVID_IMPL_FULL) {
+		drvdata->edpcsr_present  = true;
+		drvdata->edvidsr_present = true;
+	} else if (mode == EDDEVID_IMPL_EDPCSR_EDCIDSR) {
+		drvdata->edpcsr_present  = true;
+		drvdata->edvidsr_present = false;
+	} else {
+		drvdata->edpcsr_present  = false;
+		drvdata->edvidsr_present = false;
+	}
+
+	CS_LOCK(drvdata->base);
+}
+
+static int debug_probe(struct amba_device *adev, const struct amba_id *id)
+{
+	void __iomem *base;
+	struct device *dev = &adev->dev;
+	struct debug_drvdata *drvdata;
+	struct resource *res = &adev->res;
+	struct device_node *np = adev->dev.of_node;
+	char buf[32];
+	static int debug_count;
+
+	drvdata = devm_kzalloc(dev, sizeof(*drvdata), GFP_KERNEL);
+	if (!drvdata)
+		return -ENOMEM;
+
+	drvdata->cpu = np ? of_coresight_get_cpu(np) : 0;
+	drvdata->dev = &adev->dev;
+
+	dev_set_drvdata(dev, drvdata);
+
+	/* Validity for the resource is already checked by the AMBA core */
+	base = devm_ioremap_resource(dev, res);
+	if (IS_ERR(base))
+		return PTR_ERR(base);
+
+	drvdata->base = base;
+
+	get_online_cpus();
+	per_cpu(debug_drvdata, drvdata->cpu) = drvdata;
+
+	if (smp_call_function_single(drvdata->cpu,
+				debug_init_arch_data, drvdata, 1))
+		dev_err(dev, "Debug arch init failed\n");
+
+	put_online_cpus();
+
+	if (!debug_count++)
+		atomic_notifier_chain_register(&panic_notifier_list,
+					       &debug_notifier);
+
+	sprintf(buf, (char *)id->data, drvdata->cpu);
+	dev_info(dev, "%s initialized\n", buf);
+	return 0;
+}
+
+static struct amba_id debug_ids[] = {
+	{       /* Debug for Cortex-A53 */
+		.id	= 0x000bbd03,
+		.mask	= 0x000fffff,
+		.data   = "Coresight debug-CPU%d",
+	},
+	{       /* Debug for Cortex-A57 */
+		.id	= 0x000bbd07,
+		.mask	= 0x000fffff,
+		.data   = "Coresight debug-CPU%d",
+	},
+	{       /* Debug for Cortex-A72 */
+		.id	= 0x000bbd08,
+		.mask	= 0x000fffff,
+		.data   = "Coresight debug-CPU%d",
+	},
+	{ 0, 0 },
+};
+
+static struct amba_driver debug_driver = {
+	.drv = {
+		.name   = "coresight-debug",
+		.suppress_bind_attrs = true,
+	},
+	.probe		= debug_probe,
+	.id_table	= debug_ids,
+};
+builtin_amba_driver(debug_driver);
-- 
2.7.4

[toc] | [prev] | [next] | [standalone]


#1596218 — Re: [v3 3/5] coresight: add support for debug module

FromSuzuki K Poulose <Suzuki.Poulose@arm.com>
Date2017-03-09 18:10 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tj9JE-8bY-21@gated-at.bofh.it>
In reply to#1591908
On 03/03/17 06:00, Leo Yan wrote:
> Coresight includes debug module and usually the module connects with CPU
> debug logic. ARMv8 architecture reference manual (ARM DDI 0487A.k) has
> description for related info in "Part H: External Debug".
>
> Chapter H7 "The Sample-based Profiling Extension" introduces several
> sampling registers, e.g. we can check program counter value with
> combined CPU exception level, secure state, etc. So this is helpful for
> analysis CPU lockup scenarios, e.g. if one CPU has run into infinite
> loop with IRQ disabled. In this case the CPU cannot switch context and
> handle any interrupt (including IPIs), as the result it cannot handle
> SMP call for stack dump.
>
> This patch is to enable coresight debug module, so firstly this driver
> is to bind apb clock for debug module and this is to ensure the debug
> module can be accessed from program or external debugger. And the driver
> uses sample-based registers for debug purpose, e.g. when system detects
> the CPU lockup and trigger panic, the driver will dump program counter
> and combined context registers (EDCIDSR, EDVIDSR); by parsing context
> registers so can quickly get to know CPU secure state, exception level,
> etc.

The problem is, it is not guaranteed that the EDPCSR_Hi, EDCIDSR & EDVIDSR are
updated as a side effect of a memory mapped access (which is what we do here) to the
EDPCSR_Lo.

Section H.7.1.2 : Reads of EDPCSRs (in ARM DDI 0487A.k) :

"The indirect writes to EDCIDSR, EDVIDSR, and EDPCSRhi might not occur for a memory-mapped access
to the external debug interface. For more information, see Memory-mapped accesses to the external debug
interface on page H8-4968."

So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
am wondering if this is really guranteed to be useful.

>
> Some of the debug module registers are located in CPU power domain, so
> in the driver it has checked the power state for CPU before accessing
> registers within CPU power domain. For most safe way to use this driver,
> it's suggested to disable CPU low power states, this can simply set
> "nohlt" in kernel command line.
>
> Signed-off-by: Leo Yan <leo.yan@linaro.org>
> ---
>  drivers/hwtracing/coresight/Kconfig           |  10 +
>  drivers/hwtracing/coresight/Makefile          |   1 +
>  drivers/hwtracing/coresight/coresight-debug.c | 377 ++++++++++++++++++++++++++
>  3 files changed, 388 insertions(+)
>  create mode 100644 drivers/hwtracing/coresight/coresight-debug.c
>
> diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
> index 130cb21..3ed651e 100644
> --- a/drivers/hwtracing/coresight/Kconfig
> +++ b/drivers/hwtracing/coresight/Kconfig
> @@ -89,4 +89,14 @@ config CORESIGHT_STM
>  	  logging useful software events or data coming from various entities
>  	  in the system, possibly running different OSs
>
> +config CORESIGHT_DEBUG

To make it more specific, may be CORESIGHT_CPU_DEBUG ?

> +	bool "CoreSight debug driver"

"Coresight CPU Debug driver"

> +	depends on ARM || ARM64
> +	help
> +	  This driver provides support for coresight debugging module. This
> +	  is primarily used to dump sample-based profiling registers for
> +	  panic. To avoid lockups when accessing debug module registers,
> +	  it is safer to disable CPU low power states (like "nohlt" on the
> +	  kernel command line) when using this feature.
> +

> +#define EDPCSR_THUMB			BIT(0)
> +#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
> +#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)

We don't need two different masks. {ED/DBG}PCSR has only bit 0 reserved
for instruction set indication.

> +#endif
> +
> +/* bits definition for EDPRSR */
> +#define EDPRSR_DLK			BIT(6)
> +#define EDPRSR_PU			BIT(0)
> +
> +
> +static void debug_read_regs(struct debug_drvdata *drvdata)
> +{
> +	drvdata->edprsr = readl_relaxed(drvdata->base + EDPRSR);
> +
> +	if (!debug_access_permitted(drvdata))
> +		return;
> +
> +	if (!drvdata->edpcsr_present)
> +		return;
> +
> +	CS_UNLOCK(drvdata->base);
> +
> +	debug_os_unlock(drvdata);
> +
> +	drvdata->edpcsr = readl_relaxed(drvdata->base + EDPCSR);
> +
> +	/*
> +	 * As described in ARM DDI 0487A.k, if the processing
> +	 * element (PE) is in debug state, or sample-based
> +	 * profiling is prohibited, EDPCSR reads as 0xFFFFFFFF;
> +	 * EDCIDSR, EDVIDSR and EDPCSR_HI registers also become
> +	 * UNKNOWN state. So directly bail out for this case.
> +	 */
> +	if (drvdata->edpcsr == EDPCSR_PROHIBITED) {
> +		CS_LOCK(drvdata->base);
> +		return;
> +	}
> +
> +	/*
> +	 * A read of the EDPCSR normally has the side-effect of
> +	 * indirectly writing to EDCIDSR, EDVIDSR and EDPCSR_HI;
> +	 * at this point it's safe to read value from them.
> +	 */

See my comment above about the side effects of memory mapped access.

> +	drvdata->edcidsr = readl_relaxed(drvdata->base + EDCIDSR);
> +#ifdef CONFIG_64BIT
> +	drvdata->edpcsr_hi = readl_relaxed(drvdata->base + EDPCSR_HI);
> +#endif

> +
> +	if (drvdata->edvidsr_present)
> +		drvdata->edvidsr = readl_relaxed(drvdata->base + EDVIDSR);
> +
> +	CS_LOCK(drvdata->base);
> +}
> +

> +#ifndef CONFIG_64BIT

I guess this doesn't help for an ARMv8 32bit only core (e.g, Cortex-A32). And
unfortunately, there are conflicting definitions for the values for PCSROffset w.r.t
ARMv8 and ARMv7.

DBGDEVID1[3:0] For ARMv7 :

0000 - Sample offset applies based on the instruction state.
0001 - No offset applies.

EDDEVID1[3:0] For ARMv8 :
0000 - EDPCSR not implemented
0010 - EDPCSR implemented without offsets, but do not use in AArch32 state!

So there is no easy way to make sense of the value, unless you know which version
of the architecture is in use. Or may be we could co-relate it with the value from
DEVID.

i.e, EDPCSR is not implemented do not register this device, see comments on debug_probe().
( And we should also include the following test for 32bit code to see if edpcsr is implemented.
See comments on debug_init_arch_data() )


That way, we could use the following inference from the PCSROffset value :

0000 - Sample offset applies based on the instruction state (indicated by PCSR[0])
0001 - No offset applies.
0010 - No offset applies, but do not use in AArch32 mode


> +static bool debug_pc_has_offset(struct debug_drvdata *drvdata)
> +{
> +	u32 pcsr_offset;
> +
> +	pcsr_offset = drvdata->eddevid1 & EDDEVID1_PCSR_OFFSET_MASK;
> +
> +	return (pcsr_offset == EDDEVID1_PCSR_OFFSET_INS_SET);
> +}
> +
> +static unsigned long debug_adjust_pc(struct debug_drvdata *drvdata,
> +				     unsigned long pc)
> +{
> +	unsigned long arm_inst_offset = 0, thumb_inst_offset = 0;
> +
> +	if (debug_pc_has_offset(drvdata)) {
> +		arm_inst_offset = 8;
> +		thumb_inst_offset = 4;
> +	}
> +
> +	/* Handle thumb instruction */
> +	if (pc & EDPCSR_THUMB) {
> +		pc = (pc & EDPCSR_THUMB_INST_MASK) - thumb_inst_offset;
> +		return pc;
> +	}
> +
> +	/*
> +	 * Handle arm instruction offset, if the arm instruction
> +	 * is not 4 byte alignment then it's possible the case
> +	 * for implementation defined; keep original value for this
> +	 * case and print info for notice.
> +	 */
> +	if (pc & BIT(1))
> +		pr_emerg("Instruction offset is implementation defined\n");

I am struggling to find the any mention about this in the ARM ARM. Please could
you point me to it.

> +static void debug_init_arch_data(void *info)
> +{
> +	struct debug_drvdata *drvdata = info;
> +	u32 mode;
> +
> +	CS_UNLOCK(drvdata->base);
> +
> +	debug_os_unlock(drvdata);
> +
> +	/* Read device info */
> +	drvdata->eddevid  = readl_relaxed(drvdata->base + EDDEVID);
> +	drvdata->eddevid1 = readl_relaxed(drvdata->base + EDDEVID1);
> +
> +	/* Parse implementation feature */
> +	mode = drvdata->eddevid & EDDEVID_PCSAMPLE_MODE;
> +	if (mode == EDDEVID_IMPL_FULL) {
> +		drvdata->edpcsr_present  = true;
> +		drvdata->edvidsr_present = true;
> +	} else if (mode == EDDEVID_IMPL_EDPCSR_EDCIDSR) {
> +		drvdata->edpcsr_present  = true;
> +		drvdata->edvidsr_present = false;

As discussed above, we need to consult the DEVID1:PCSROffset for AArch32 to decide
if we have the edpcsr implemented on ARMv8.

> +	} else {
> +		drvdata->edpcsr_present  = false;
> +		drvdata->edvidsr_present = false;
> +	}
> +
> +	CS_LOCK(drvdata->base);
> +}
> +
> +static int debug_probe(struct amba_device *adev, const struct amba_id *id)
> +{
> +	void __iomem *base;
> +	struct device *dev = &adev->dev;
> +	struct debug_drvdata *drvdata;
> +	struct resource *res = &adev->res;
> +	struct device_node *np = adev->dev.of_node;
> +	char buf[32];
> +	static int debug_count;
> +
> +	drvdata = devm_kzalloc(dev, sizeof(*drvdata), GFP_KERNEL);
> +	if (!drvdata)
> +		return -ENOMEM;
> +
> +	drvdata->cpu = np ? of_coresight_get_cpu(np) : 0;
> +	drvdata->dev = &adev->dev;
> +
> +	dev_set_drvdata(dev, drvdata);
> +
> +	/* Validity for the resource is already checked by the AMBA core */
> +	base = devm_ioremap_resource(dev, res);
> +	if (IS_ERR(base))
> +		return PTR_ERR(base);
> +
> +	drvdata->base = base;
> +
> +	get_online_cpus();
> +	per_cpu(debug_drvdata, drvdata->cpu) = drvdata;
> +
> +	if (smp_call_function_single(drvdata->cpu,
> +				debug_init_arch_data, drvdata, 1))
> +		dev_err(dev, "Debug arch init failed\n");

If this fails (say the CPU was offline), should we still return success ?
And may be we should check if the drvdata->edpcsr_present to detect if the CPU
implements the PC Sampling and return failure here if it doesn't.

> +
> +	put_online_cpus();
> +
> +	if (!debug_count++)
> +		atomic_notifier_chain_register(&panic_notifier_list,
> +					       &debug_notifier);
> +

> +	sprintf(buf, (char *)id->data, drvdata->cpu);
> +	dev_info(dev, "%s initialized\n", buf);

This could simply be :
	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);

and get rid of the static string and the buffer, see below.

> +	return 0;
> +}
> +
> +static struct amba_id debug_ids[] = {
> +	{       /* Debug for Cortex-A53 */
> +		.id	= 0x000bbd03,
> +		.mask	= 0x000fffff,

...

> +		.data   = "Coresight debug-CPU%d",

I think this is pointless, as the debug area we are interested in is always associated
with a CPU, we could as well figure out what to print from the drvdata->cpu above.

Suzuki

[toc] | [prev] | [next] | [standalone]


#1596281 — Re: [v3 3/5] coresight: add support for debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-09 19:10 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tjaFI-ma-31@gated-at.bofh.it>
In reply to#1596218
Hi Suziku,

Thanks for reviewing, please see some replying.

On Thu, Mar 09, 2017 at 04:53:05PM +0000, Suzuki K Poulose wrote:
> On 03/03/17 06:00, Leo Yan wrote:
> >Coresight includes debug module and usually the module connects with CPU
> >debug logic. ARMv8 architecture reference manual (ARM DDI 0487A.k) has
> >description for related info in "Part H: External Debug".
> >
> >Chapter H7 "The Sample-based Profiling Extension" introduces several
> >sampling registers, e.g. we can check program counter value with
> >combined CPU exception level, secure state, etc. So this is helpful for
> >analysis CPU lockup scenarios, e.g. if one CPU has run into infinite
> >loop with IRQ disabled. In this case the CPU cannot switch context and
> >handle any interrupt (including IPIs), as the result it cannot handle
> >SMP call for stack dump.
> >
> >This patch is to enable coresight debug module, so firstly this driver
> >is to bind apb clock for debug module and this is to ensure the debug
> >module can be accessed from program or external debugger. And the driver
> >uses sample-based registers for debug purpose, e.g. when system detects
> >the CPU lockup and trigger panic, the driver will dump program counter
> >and combined context registers (EDCIDSR, EDVIDSR); by parsing context
> >registers so can quickly get to know CPU secure state, exception level,
> >etc.
> 
> The problem is, it is not guaranteed that the EDPCSR_Hi, EDCIDSR & EDVIDSR are
> updated as a side effect of a memory mapped access (which is what we do here) to the
> EDPCSR_Lo.
> 
> Section H.7.1.2 : Reads of EDPCSRs (in ARM DDI 0487A.k) :
> 
> "The indirect writes to EDCIDSR, EDVIDSR, and EDPCSRhi might not occur for a memory-mapped access
> to the external debug interface. For more information, see Memory-mapped accesses to the external debug
> interface on page H8-4968."
> 
> So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
> am wondering if this is really guranteed to be useful.

So this is caused by Software lock is locked?

Section H8.4.1: 

"Reads and writes have no side-effects. A side-effect is where a
direct read or a direct write of a register creates
an indirect write of the same or another register. When the Software
Lock is locked, the indirect write does
not occur."

> >Some of the debug module registers are located in CPU power domain, so
> >in the driver it has checked the power state for CPU before accessing
> >registers within CPU power domain. For most safe way to use this driver,
> >it's suggested to disable CPU low power states, this can simply set
> >"nohlt" in kernel command line.
> >
> >Signed-off-by: Leo Yan <leo.yan@linaro.org>
> >---
> > drivers/hwtracing/coresight/Kconfig           |  10 +
> > drivers/hwtracing/coresight/Makefile          |   1 +
> > drivers/hwtracing/coresight/coresight-debug.c | 377 ++++++++++++++++++++++++++
> > 3 files changed, 388 insertions(+)
> > create mode 100644 drivers/hwtracing/coresight/coresight-debug.c
> >
> >diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
> >index 130cb21..3ed651e 100644
> >--- a/drivers/hwtracing/coresight/Kconfig
> >+++ b/drivers/hwtracing/coresight/Kconfig
> >@@ -89,4 +89,14 @@ config CORESIGHT_STM
> > 	  logging useful software events or data coming from various entities
> > 	  in the system, possibly running different OSs
> >
> >+config CORESIGHT_DEBUG
> 
> To make it more specific, may be CORESIGHT_CPU_DEBUG ?

Will fix.

> >+	bool "CoreSight debug driver"
> 
> "Coresight CPU Debug driver"

Will fix.

> >+	depends on ARM || ARM64
> >+	help
> >+	  This driver provides support for coresight debugging module. This
> >+	  is primarily used to dump sample-based profiling registers for
> >+	  panic. To avoid lockups when accessing debug module registers,
> >+	  it is safer to disable CPU low power states (like "nohlt" on the
> >+	  kernel command line) when using this feature.
> >+
> 
> >+#define EDPCSR_THUMB			BIT(0)
> >+#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
> >+#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)
> 
> We don't need two different masks. {ED/DBG}PCSR has only bit 0 reserved
> for instruction set indication.

I think we need this two different masks. Please review below extra doc
for PC offset analysis in ARM ARM.

> >+#endif
> >+
> >+/* bits definition for EDPRSR */
> >+#define EDPRSR_DLK			BIT(6)
> >+#define EDPRSR_PU			BIT(0)
> >+
> >+
> >+static void debug_read_regs(struct debug_drvdata *drvdata)
> >+{
> >+	drvdata->edprsr = readl_relaxed(drvdata->base + EDPRSR);
> >+
> >+	if (!debug_access_permitted(drvdata))
> >+		return;
> >+
> >+	if (!drvdata->edpcsr_present)
> >+		return;
> >+
> >+	CS_UNLOCK(drvdata->base);
> >+
> >+	debug_os_unlock(drvdata);
> >+
> >+	drvdata->edpcsr = readl_relaxed(drvdata->base + EDPCSR);
> >+
> >+	/*
> >+	 * As described in ARM DDI 0487A.k, if the processing
> >+	 * element (PE) is in debug state, or sample-based
> >+	 * profiling is prohibited, EDPCSR reads as 0xFFFFFFFF;
> >+	 * EDCIDSR, EDVIDSR and EDPCSR_HI registers also become
> >+	 * UNKNOWN state. So directly bail out for this case.
> >+	 */
> >+	if (drvdata->edpcsr == EDPCSR_PROHIBITED) {
> >+		CS_LOCK(drvdata->base);
> >+		return;
> >+	}
> >+
> >+	/*
> >+	 * A read of the EDPCSR normally has the side-effect of
> >+	 * indirectly writing to EDCIDSR, EDVIDSR and EDPCSR_HI;
> >+	 * at this point it's safe to read value from them.
> >+	 */
> 
> See my comment above about the side effects of memory mapped access.

Yeah.

> >+	drvdata->edcidsr = readl_relaxed(drvdata->base + EDCIDSR);
> >+#ifdef CONFIG_64BIT
> >+	drvdata->edpcsr_hi = readl_relaxed(drvdata->base + EDPCSR_HI);
> >+#endif
> 
> >+
> >+	if (drvdata->edvidsr_present)
> >+		drvdata->edvidsr = readl_relaxed(drvdata->base + EDVIDSR);
> >+
> >+	CS_LOCK(drvdata->base);
> >+}
> >+
> 
> >+#ifndef CONFIG_64BIT
> 
> I guess this doesn't help for an ARMv8 32bit only core (e.g, Cortex-A32). And
> unfortunately, there are conflicting definitions for the values for PCSROffset w.r.t
> ARMv8 and ARMv7.
> 
> DBGDEVID1[3:0] For ARMv7 :
> 
> 0000 - Sample offset applies based on the instruction state.
> 0001 - No offset applies.
> 
> EDDEVID1[3:0] For ARMv8 :
> 0000 - EDPCSR not implemented
> 0010 - EDPCSR implemented without offsets, but do not use in AArch32 state!
> 
> So there is no easy way to make sense of the value, unless you know which version
> of the architecture is in use. Or may be we could co-relate it with the value from
> DEVID.
> 
> i.e, EDPCSR is not implemented do not register this device, see comments on debug_probe().
> ( And we should also include the following test for 32bit code to see if edpcsr is implemented.
> See comments on debug_init_arch_data() )
> 
> 
> That way, we could use the following inference from the PCSROffset value :
> 
> 0000 - Sample offset applies based on the instruction state (indicated by PCSR[0])
> 0001 - No offset applies.
> 0010 - No offset applies, but do not use in AArch32 mode

Just now I went through ARM ARM and ARMv8 ARM, this makes sense to me.
Thanks for good pointing for this.

> >+static bool debug_pc_has_offset(struct debug_drvdata *drvdata)
> >+{
> >+	u32 pcsr_offset;
> >+
> >+	pcsr_offset = drvdata->eddevid1 & EDDEVID1_PCSR_OFFSET_MASK;
> >+
> >+	return (pcsr_offset == EDDEVID1_PCSR_OFFSET_INS_SET);
> >+}
> >+
> >+static unsigned long debug_adjust_pc(struct debug_drvdata *drvdata,
> >+				     unsigned long pc)
> >+{
> >+	unsigned long arm_inst_offset = 0, thumb_inst_offset = 0;
> >+
> >+	if (debug_pc_has_offset(drvdata)) {
> >+		arm_inst_offset = 8;
> >+		thumb_inst_offset = 4;
> >+	}
> >+
> >+	/* Handle thumb instruction */
> >+	if (pc & EDPCSR_THUMB) {
> >+		pc = (pc & EDPCSR_THUMB_INST_MASK) - thumb_inst_offset;
> >+		return pc;
> >+	}
> >+
> >+	/*
> >+	 * Handle arm instruction offset, if the arm instruction
> >+	 * is not 4 byte alignment then it's possible the case
> >+	 * for implementation defined; keep original value for this
> >+	 * case and print info for notice.
> >+	 */
> >+	if (pc & BIT(1))
> >+		pr_emerg("Instruction offset is implementation defined\n");
> 
> I am struggling to find the any mention about this in the ARM ARM. Please could
> you point me to it.

Sure, please see ARM DDI 0406C.b, chapter C11.11.34 "
DBGPCSR, Program Counter Sampling Register":

A profiling tool can use the value of the T bit to calculate the
instruction address as follows:

When an offset is applied to the sampled address
- if T is 0 and DBGPCSR[1] is 0, ((DBGPCSR[31:2] << 2) - 8) is the
address of the sampled ARM instruction
- if T is 0 and DBGPCSR[1] is 1, the instruction address is
IMPLEMENTATION DEFINED
- if T is 1, ((DBGPCSR[31:1] << 1) - 4) is the address of the sampled
Thumb or ThumbEE instruction.

When no offset is applied to the sampled address
-  if T is 0 and DBGPCSR[1] is 0, (DBGPCSR[31:2] << 2) is the address
of the sampled ARM instruction
-  if T is 0 and DBGPCSR[1] is 1, the instruction address is
IMPLEMENTATION DEFINED
- if T is 1, (DBGPCSR[31:1] << 1) is the address of the sampled Thumb
or ThumbEE instruction.

> >+static void debug_init_arch_data(void *info)
> >+{
> >+	struct debug_drvdata *drvdata = info;
> >+	u32 mode;
> >+
> >+	CS_UNLOCK(drvdata->base);
> >+
> >+	debug_os_unlock(drvdata);
> >+
> >+	/* Read device info */
> >+	drvdata->eddevid  = readl_relaxed(drvdata->base + EDDEVID);
> >+	drvdata->eddevid1 = readl_relaxed(drvdata->base + EDDEVID1);
> >+
> >+	/* Parse implementation feature */
> >+	mode = drvdata->eddevid & EDDEVID_PCSAMPLE_MODE;
> >+	if (mode == EDDEVID_IMPL_FULL) {
> >+		drvdata->edpcsr_present  = true;
> >+		drvdata->edvidsr_present = true;
> >+	} else if (mode == EDDEVID_IMPL_EDPCSR_EDCIDSR) {
> >+		drvdata->edpcsr_present  = true;
> >+		drvdata->edvidsr_present = false;
> 
> As discussed above, we need to consult the DEVID1:PCSROffset for AArch32 to decide
> if we have the edpcsr implemented on ARMv8.

Yeah.

> >+	} else {
> >+		drvdata->edpcsr_present  = false;
> >+		drvdata->edvidsr_present = false;
> >+	}
> >+
> >+	CS_LOCK(drvdata->base);
> >+}
> >+
> >+static int debug_probe(struct amba_device *adev, const struct amba_id *id)
> >+{
> >+	void __iomem *base;
> >+	struct device *dev = &adev->dev;
> >+	struct debug_drvdata *drvdata;
> >+	struct resource *res = &adev->res;
> >+	struct device_node *np = adev->dev.of_node;
> >+	char buf[32];
> >+	static int debug_count;
> >+
> >+	drvdata = devm_kzalloc(dev, sizeof(*drvdata), GFP_KERNEL);
> >+	if (!drvdata)
> >+		return -ENOMEM;
> >+
> >+	drvdata->cpu = np ? of_coresight_get_cpu(np) : 0;
> >+	drvdata->dev = &adev->dev;
> >+
> >+	dev_set_drvdata(dev, drvdata);
> >+
> >+	/* Validity for the resource is already checked by the AMBA core */
> >+	base = devm_ioremap_resource(dev, res);
> >+	if (IS_ERR(base))
> >+		return PTR_ERR(base);
> >+
> >+	drvdata->base = base;
> >+
> >+	get_online_cpus();
> >+	per_cpu(debug_drvdata, drvdata->cpu) = drvdata;
> >+
> >+	if (smp_call_function_single(drvdata->cpu,
> >+				debug_init_arch_data, drvdata, 1))
> >+		dev_err(dev, "Debug arch init failed\n");
> 
> If this fails (say the CPU was offline), should we still return success ?
> And may be we should check if the drvdata->edpcsr_present to detect if the CPU
> implements the PC Sampling and return failure here if it doesn't.

Will fix.

> >+
> >+	put_online_cpus();
> >+
> >+	if (!debug_count++)
> >+		atomic_notifier_chain_register(&panic_notifier_list,
> >+					       &debug_notifier);
> >+
> 
> >+	sprintf(buf, (char *)id->data, drvdata->cpu);
> >+	dev_info(dev, "%s initialized\n", buf);
> 
> This could simply be :
> 	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);
> 
> and get rid of the static string and the buffer, see below.
> 
> >+	return 0;
> >+}
> >+
> >+static struct amba_id debug_ids[] = {
> >+	{       /* Debug for Cortex-A53 */
> >+		.id	= 0x000bbd03,
> >+		.mask	= 0x000fffff,
> 
> ...
> 
> >+		.data   = "Coresight debug-CPU%d",
> 
> I think this is pointless, as the debug area we are interested in is always associated
> with a CPU, we could as well figure out what to print from the drvdata->cpu above.

I prefer to follow your suggestion for upper two comments; but I'd like
check with Mathieu, due I followed up Mathieu's suggestion to write
current code.

Thanks,
Leo Yan

[toc] | [prev] | [next] | [standalone]


#1597866 — Re: [v3 3/5] coresight: add support for debug module

FromSuzuki K Poulose <Suzuki.Poulose@arm.com>
Date2017-03-10 15:40 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tjtS3-5k0-39@gated-at.bofh.it>
In reply to#1596281
On 09/03/17 17:59, Leo Yan wrote:
> Hi Suziku,
  
>> The problem is, it is not guaranteed that the EDPCSR_Hi, EDCIDSR & EDVIDSR are
>> updated as a side effect of a memory mapped access (which is what we do here) to the
>> EDPCSR_Lo.
>>
>> Section H.7.1.2 : Reads of EDPCSRs (in ARM DDI 0487A.k) :
>>
>> "The indirect writes to EDCIDSR, EDVIDSR, and EDPCSRhi might not occur for a memory-mapped access
>> to the external debug interface. For more information, see Memory-mapped accesses to the external debug
>> interface on page H8-4968."
>>
>> So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
>> am wondering if this is really guranteed to be useful.
>
> So this is caused by Software lock is locked?
>
> Section H8.4.1:
>
> "Reads and writes have no side-effects. A side-effect is where a
> direct read or a direct write of a register creates
> an indirect write of the same or another register. When the Software
> Lock is locked, the indirect write does
> not occur."

Yes, thats correct, further :

Section H9.2.32: EDPCSR

"For a read of EDPCSRlo from the memory-mapped interface, if EDLSR.SLK == 1, meaning
the Software Lock is locked, then the access has no side-effects. That is, EDCIDSR,
EDVIDSR, and EDPCSRhi are unchanged."

And since we do a CS_UNLOCK, that should be fine. Please ignore my comment.

>>> diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
>>> index 130cb21..3ed651e 100644
>>> --- a/drivers/hwtracing/coresight/Kconfig
>>> +++ b/drivers/hwtracing/coresight/Kconfig
>>> @@ -89,4 +89,14 @@ config CORESIGHT_STM
>>> 	  logging useful software events or data coming from various entities
>>> 	  in the system, possibly running different OSs
>>>
>>> +config CORESIGHT_DEBUG
>>
>> To make it more specific, may be CORESIGHT_CPU_DEBUG ?
>
> Will fix.
>
>>> +	bool "CoreSight debug driver"
>>
>> "Coresight CPU Debug driver"
>
> Will fix.
>
>>> +	depends on ARM || ARM64
>>> +	help
>>> +	  This driver provides support for coresight debugging module. This
>>> +	  is primarily used to dump sample-based profiling registers for
>>> +	  panic. To avoid lockups when accessing debug module registers,
>>> +	  it is safer to disable CPU low power states (like "nohlt" on the
>>> +	  kernel command line) when using this feature.
>>> +
>>
>>> +#define EDPCSR_THUMB			BIT(0)
>>> +#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
>>> +#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)
>>
>> We don't need two different masks. {ED/DBG}PCSR has only bit 0 reserved
>> for instruction set indication.
>
> I think we need this two different masks. Please review below extra doc
> for PC offset analysis in ARM ARM.

You're correct. Thanks for the pointer. I got confused, as there was no bit
dedicated in the DBGPCSR bit assignment figure.

>>> +	/*
>>> +	 * Handle arm instruction offset, if the arm instruction
>>> +	 * is not 4 byte alignment then it's possible the case
>>> +	 * for implementation defined; keep original value for this
>>> +	 * case and print info for notice.
>>> +	 */
>>> +	if (pc & BIT(1))
>>> +		pr_emerg("Instruction offset is implementation defined\n");
>>
>> I am struggling to find the any mention about this in the ARM ARM. Please could
>> you point me to it.
>
> Sure, please see ARM DDI 0406C.b, chapter C11.11.34 "
> DBGPCSR, Program Counter Sampling Register":
>
> A profiling tool can use the value of the T bit to calculate the
> instruction address as follows:
>
> When an offset is applied to the sampled address
> - if T is 0 and DBGPCSR[1] is 0, ((DBGPCSR[31:2] << 2) - 8) is the
> address of the sampled ARM instruction
> - if T is 0 and DBGPCSR[1] is 1, the instruction address is
> IMPLEMENTATION DEFINED
> - if T is 1, ((DBGPCSR[31:1] << 1) - 4) is the address of the sampled
> Thumb or ThumbEE instruction.
>
> When no offset is applied to the sampled address
> -  if T is 0 and DBGPCSR[1] is 0, (DBGPCSR[31:2] << 2) is the address
> of the sampled ARM instruction
> -  if T is 0 and DBGPCSR[1] is 1, the instruction address is
> IMPLEMENTATION DEFINED
> - if T is 1, (DBGPCSR[31:1] << 1) is the address of the sampled Thumb
> or ThumbEE instruction.
>

Ok.


>>> +
>>> +	put_online_cpus();
>>> +
>>> +	if (!debug_count++)
>>> +		atomic_notifier_chain_register(&panic_notifier_list,
>>> +					       &debug_notifier);
>>> +
>>
>>> +	sprintf(buf, (char *)id->data, drvdata->cpu);
>>> +	dev_info(dev, "%s initialized\n", buf);
>>
>> This could simply be :
>> 	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);
>>
>> and get rid of the static string and the buffer, see below.

Also we need pm_runtime_put() here to balance the pm_runtime_get_ from AMBA
device probe. More on that below.

>>
>>> +	return 0;
>>> +}
>>> +
>>> +static struct amba_id debug_ids[] = {
>>> +	{       /* Debug for Cortex-A53 */
>>> +		.id	= 0x000bbd03,
>>> +		.mask	= 0x000fffff,
>>
>> ...
>>
>>> +		.data   = "Coresight debug-CPU%d",
>>
>> I think this is pointless, as the debug area we are interested in is always associated
>> with a CPU, we could as well figure out what to print from the drvdata->cpu above.
>
> I prefer to follow your suggestion for upper two comments; but I'd like
> check with Mathieu, due I followed up Mathieu's suggestion to write
> current code.

Btw, I don't see any PM calls to make sure the power domain (at least the debug domain)
is up, which could cause problems with accesses to some of these registers (leave alone the
ones in CPU power domain), especially the EDPRSR. We could also do pm_runtime_get on the
CPU's power domain, if the CPU is online, before we access the pcsr.

Suzuki

[toc] | [prev] | [next] | [standalone]


#1598931 — Re: [v3 3/5] coresight: add support for debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-13 09:20 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tktmW-66Q-15@gated-at.bofh.it>
In reply to#1597866
Hi Suzuki,

On Fri, Mar 10, 2017 at 02:29:53PM +0000, Suzuki K Poulose wrote:

[...]

> >>So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
> >>am wondering if this is really guranteed to be useful.
> >
> >So this is caused by Software lock is locked?
> >
> >Section H8.4.1:
> >
> >"Reads and writes have no side-effects. A side-effect is where a
> >direct read or a direct write of a register creates
> >an indirect write of the same or another register. When the Software
> >Lock is locked, the indirect write does
> >not occur."
> 
> Yes, thats correct, further :
> 
> Section H9.2.32: EDPCSR
> 
> "For a read of EDPCSRlo from the memory-mapped interface, if EDLSR.SLK == 1, meaning
> the Software Lock is locked, then the access has no side-effects. That is, EDCIDSR,
> EDVIDSR, and EDPCSRhi are unchanged."
> 
> And since we do a CS_UNLOCK, that should be fine. Please ignore my comment.

Thanks a lot for confirmation.

[...]

> >>>+
> >>>+	put_online_cpus();
> >>>+
> >>>+	if (!debug_count++)
> >>>+		atomic_notifier_chain_register(&panic_notifier_list,
> >>>+					       &debug_notifier);
> >>>+
> >>
> >>>+	sprintf(buf, (char *)id->data, drvdata->cpu);
> >>>+	dev_info(dev, "%s initialized\n", buf);
> >>
> >>This could simply be :
> >>	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);
> >>
> >>and get rid of the static string and the buffer, see below.
> 
> Also we need pm_runtime_put() here to balance the pm_runtime_get_ from AMBA
> device probe. More on that below.

[...]

> Btw, I don't see any PM calls to make sure the power domain (at least the debug domain)
> is up, which could cause problems with accesses to some of these registers (leave alone the
> ones in CPU power domain), especially the EDPRSR. We could also do pm_runtime_get on the
> CPU's power domain, if the CPU is online, before we access the pcsr.

I will add pm_runtime_get/pm_runtime_put for apb clock.

But for CPU power domain, AFAIK this part is managed by PSCI but is not
controlled by pm_runtime_{put|get} pairs. So at beginning, we suggest
to use "nohlt" to ensure CPU power domain is enabled.

Please let me know if I miss some thing for this?

Thanks,
Leo Yan

[toc] | [prev] | [next] | [standalone]


#1599625 — Re: [v3 3/5] coresight: add support for debug module

FromMathieu Poirier <mathieu.poirier@linaro.org>
Date2017-03-13 18:00 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tkBub-3kO-37@gated-at.bofh.it>
In reply to#1597866
On Fri, Mar 10, 2017 at 02:29:53PM +0000, Suzuki K Poulose wrote:
> On 09/03/17 17:59, Leo Yan wrote:
> >Hi Suziku,
> >>The problem is, it is not guaranteed that the EDPCSR_Hi, EDCIDSR & EDVIDSR are
> >>updated as a side effect of a memory mapped access (which is what we do here) to the
> >>EDPCSR_Lo.
> >>
> >>Section H.7.1.2 : Reads of EDPCSRs (in ARM DDI 0487A.k) :
> >>
> >>"The indirect writes to EDCIDSR, EDVIDSR, and EDPCSRhi might not occur for a memory-mapped access
> >>to the external debug interface. For more information, see Memory-mapped accesses to the external debug
> >>interface on page H8-4968."
> >>
> >>So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
> >>am wondering if this is really guranteed to be useful.
> >
> >So this is caused by Software lock is locked?
> >
> >Section H8.4.1:
> >
> >"Reads and writes have no side-effects. A side-effect is where a
> >direct read or a direct write of a register creates
> >an indirect write of the same or another register. When the Software
> >Lock is locked, the indirect write does
> >not occur."
> 
> Yes, thats correct, further :
> 
> Section H9.2.32: EDPCSR
> 
> "For a read of EDPCSRlo from the memory-mapped interface, if EDLSR.SLK == 1, meaning
> the Software Lock is locked, then the access has no side-effects. That is, EDCIDSR,
> EDVIDSR, and EDPCSRhi are unchanged."
> 
> And since we do a CS_UNLOCK, that should be fine. Please ignore my comment.
> 
> >>>diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
> >>>index 130cb21..3ed651e 100644
> >>>--- a/drivers/hwtracing/coresight/Kconfig
> >>>+++ b/drivers/hwtracing/coresight/Kconfig
> >>>@@ -89,4 +89,14 @@ config CORESIGHT_STM
> >>>	  logging useful software events or data coming from various entities
> >>>	  in the system, possibly running different OSs
> >>>
> >>>+config CORESIGHT_DEBUG
> >>
> >>To make it more specific, may be CORESIGHT_CPU_DEBUG ?
> >
> >Will fix.
> >
> >>>+	bool "CoreSight debug driver"
> >>
> >>"Coresight CPU Debug driver"
> >
> >Will fix.
> >
> >>>+	depends on ARM || ARM64
> >>>+	help
> >>>+	  This driver provides support for coresight debugging module. This
> >>>+	  is primarily used to dump sample-based profiling registers for
> >>>+	  panic. To avoid lockups when accessing debug module registers,
> >>>+	  it is safer to disable CPU low power states (like "nohlt" on the
> >>>+	  kernel command line) when using this feature.
> >>>+
> >>
> >>>+#define EDPCSR_THUMB			BIT(0)
> >>>+#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
> >>>+#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)
> >>
> >>We don't need two different masks. {ED/DBG}PCSR has only bit 0 reserved
> >>for instruction set indication.
> >
> >I think we need this two different masks. Please review below extra doc
> >for PC offset analysis in ARM ARM.
> 
> You're correct. Thanks for the pointer. I got confused, as there was no bit
> dedicated in the DBGPCSR bit assignment figure.
> 
> >>>+	/*
> >>>+	 * Handle arm instruction offset, if the arm instruction
> >>>+	 * is not 4 byte alignment then it's possible the case
> >>>+	 * for implementation defined; keep original value for this
> >>>+	 * case and print info for notice.
> >>>+	 */
> >>>+	if (pc & BIT(1))
> >>>+		pr_emerg("Instruction offset is implementation defined\n");
> >>
> >>I am struggling to find the any mention about this in the ARM ARM. Please could
> >>you point me to it.
> >
> >Sure, please see ARM DDI 0406C.b, chapter C11.11.34 "
> >DBGPCSR, Program Counter Sampling Register":
> >
> >A profiling tool can use the value of the T bit to calculate the
> >instruction address as follows:
> >
> >When an offset is applied to the sampled address
> >- if T is 0 and DBGPCSR[1] is 0, ((DBGPCSR[31:2] << 2) - 8) is the
> >address of the sampled ARM instruction
> >- if T is 0 and DBGPCSR[1] is 1, the instruction address is
> >IMPLEMENTATION DEFINED
> >- if T is 1, ((DBGPCSR[31:1] << 1) - 4) is the address of the sampled
> >Thumb or ThumbEE instruction.
> >
> >When no offset is applied to the sampled address
> >-  if T is 0 and DBGPCSR[1] is 0, (DBGPCSR[31:2] << 2) is the address
> >of the sampled ARM instruction
> >-  if T is 0 and DBGPCSR[1] is 1, the instruction address is
> >IMPLEMENTATION DEFINED
> >- if T is 1, (DBGPCSR[31:1] << 1) is the address of the sampled Thumb
> >or ThumbEE instruction.
> >
> 
> Ok.
> 
> 
> >>>+
> >>>+	put_online_cpus();
> >>>+
> >>>+	if (!debug_count++)
> >>>+		atomic_notifier_chain_register(&panic_notifier_list,
> >>>+					       &debug_notifier);
> >>>+
> >>
> >>>+	sprintf(buf, (char *)id->data, drvdata->cpu);
> >>>+	dev_info(dev, "%s initialized\n", buf);
> >>
> >>This could simply be :
> >>	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);
> >>
> >>and get rid of the static string and the buffer, see below.
> 
> Also we need pm_runtime_put() here to balance the pm_runtime_get_ from AMBA
> device probe.

Good point.

> More on that below.
> 
> >>
> >>>+	return 0;
> >>>+}
> >>>+
> >>>+static struct amba_id debug_ids[] = {
> >>>+	{       /* Debug for Cortex-A53 */
> >>>+		.id	= 0x000bbd03,
> >>>+		.mask	= 0x000fffff,
> >>
> >>...
> >>
> >>>+		.data   = "Coresight debug-CPU%d",
> >>
> >>I think this is pointless, as the debug area we are interested in is always associated
> >>with a CPU, we could as well figure out what to print from the drvdata->cpu above.
> >
> >I prefer to follow your suggestion for upper two comments; but I'd like
> >check with Mathieu, due I followed up Mathieu's suggestion to write
> >current code.
> 
> Btw, I don't see any PM calls to make sure the power domain (at least the debug domain)
> is up, which could cause problems with accesses to some of these registers (leave alone the
> ones in CPU power domain), especially the EDPRSR. We could also do pm_runtime_get on the
> CPU's power domain, if the CPU is online, before we access the pcsr.

I thought about PM runtime operations a little while back but wondered if it is
really a good thing to have them around.  When this code is called the system
has crashed and as such making PM runtimes call isn't a good idea.

One thing we could do is _not_ call pm_runtime_put() at the end of the probe()
operation.  That way we wouldn't have to mess around with PM runtime operations
on an unstable system.  This, of course, is costly in terms of power consumption
but the system is under test/debug anyway.

Thoughts?

> 
> Suzuki
> 
> 
> 
> 
> 

[toc] | [prev] | [next] | [standalone]


#1599579 — Re: [v3 3/5] coresight: add support for debug module

FromMathieu Poirier <mathieu.poirier@linaro.org>
Date2017-03-13 17:40 +0100
SubjectRe: [v3 3/5] coresight: add support for debug module
Message-ID<tkBaO-3bi-21@gated-at.bofh.it>
In reply to#1596281
On Fri, Mar 10, 2017 at 01:59:15AM +0800, Leo Yan wrote:
> Hi Suziku,
> 
> Thanks for reviewing, please see some replying.
> 
> On Thu, Mar 09, 2017 at 04:53:05PM +0000, Suzuki K Poulose wrote:
> > On 03/03/17 06:00, Leo Yan wrote:
> > >Coresight includes debug module and usually the module connects with CPU
> > >debug logic. ARMv8 architecture reference manual (ARM DDI 0487A.k) has
> > >description for related info in "Part H: External Debug".
> > >
> > >Chapter H7 "The Sample-based Profiling Extension" introduces several
> > >sampling registers, e.g. we can check program counter value with
> > >combined CPU exception level, secure state, etc. So this is helpful for
> > >analysis CPU lockup scenarios, e.g. if one CPU has run into infinite
> > >loop with IRQ disabled. In this case the CPU cannot switch context and
> > >handle any interrupt (including IPIs), as the result it cannot handle
> > >SMP call for stack dump.
> > >
> > >This patch is to enable coresight debug module, so firstly this driver
> > >is to bind apb clock for debug module and this is to ensure the debug
> > >module can be accessed from program or external debugger. And the driver
> > >uses sample-based registers for debug purpose, e.g. when system detects
> > >the CPU lockup and trigger panic, the driver will dump program counter
> > >and combined context registers (EDCIDSR, EDVIDSR); by parsing context
> > >registers so can quickly get to know CPU secure state, exception level,
> > >etc.
> > 
> > The problem is, it is not guaranteed that the EDPCSR_Hi, EDCIDSR & EDVIDSR are
> > updated as a side effect of a memory mapped access (which is what we do here) to the
> > EDPCSR_Lo.
> > 
> > Section H.7.1.2 : Reads of EDPCSRs (in ARM DDI 0487A.k) :
> > 
> > "The indirect writes to EDCIDSR, EDVIDSR, and EDPCSRhi might not occur for a memory-mapped access
> > to the external debug interface. For more information, see Memory-mapped accesses to the external debug
> > interface on page H8-4968."
> > 
> > So we cannot really rely on the values in EDVIDSR which we use to make further decisions. So I
> > am wondering if this is really guranteed to be useful.
> 
> So this is caused by Software lock is locked?
> 
> Section H8.4.1: 
> 
> "Reads and writes have no side-effects. A side-effect is where a
> direct read or a direct write of a register creates
> an indirect write of the same or another register. When the Software
> Lock is locked, the indirect write does
> not occur."
> 
> > >Some of the debug module registers are located in CPU power domain, so
> > >in the driver it has checked the power state for CPU before accessing
> > >registers within CPU power domain. For most safe way to use this driver,
> > >it's suggested to disable CPU low power states, this can simply set
> > >"nohlt" in kernel command line.
> > >
> > >Signed-off-by: Leo Yan <leo.yan@linaro.org>
> > >---
> > > drivers/hwtracing/coresight/Kconfig           |  10 +
> > > drivers/hwtracing/coresight/Makefile          |   1 +
> > > drivers/hwtracing/coresight/coresight-debug.c | 377 ++++++++++++++++++++++++++
> > > 3 files changed, 388 insertions(+)
> > > create mode 100644 drivers/hwtracing/coresight/coresight-debug.c
> > >
> > >diff --git a/drivers/hwtracing/coresight/Kconfig b/drivers/hwtracing/coresight/Kconfig
> > >index 130cb21..3ed651e 100644
> > >--- a/drivers/hwtracing/coresight/Kconfig
> > >+++ b/drivers/hwtracing/coresight/Kconfig
> > >@@ -89,4 +89,14 @@ config CORESIGHT_STM
> > > 	  logging useful software events or data coming from various entities
> > > 	  in the system, possibly running different OSs
> > >
> > >+config CORESIGHT_DEBUG
> > 
> > To make it more specific, may be CORESIGHT_CPU_DEBUG ?
> 
> Will fix.
> 
> > >+	bool "CoreSight debug driver"
> > 
> > "Coresight CPU Debug driver"
> 
> Will fix.
> 
> > >+	depends on ARM || ARM64
> > >+	help
> > >+	  This driver provides support for coresight debugging module. This
> > >+	  is primarily used to dump sample-based profiling registers for
> > >+	  panic. To avoid lockups when accessing debug module registers,
> > >+	  it is safer to disable CPU low power states (like "nohlt" on the
> > >+	  kernel command line) when using this feature.
> > >+
> > 
> > >+#define EDPCSR_THUMB			BIT(0)
> > >+#define EDPCSR_ARM_INST_MASK		GENMASK(31, 2)
> > >+#define EDPCSR_THUMB_INST_MASK		GENMASK(31, 1)
> > 
> > We don't need two different masks. {ED/DBG}PCSR has only bit 0 reserved
> > for instruction set indication.
> 
> I think we need this two different masks. Please review below extra doc
> for PC offset analysis in ARM ARM.
> 
> > >+#endif
> > >+
> > >+/* bits definition for EDPRSR */
> > >+#define EDPRSR_DLK			BIT(6)
> > >+#define EDPRSR_PU			BIT(0)
> > >+
> > >+
> > >+static void debug_read_regs(struct debug_drvdata *drvdata)
> > >+{
> > >+	drvdata->edprsr = readl_relaxed(drvdata->base + EDPRSR);
> > >+
> > >+	if (!debug_access_permitted(drvdata))
> > >+		return;
> > >+
> > >+	if (!drvdata->edpcsr_present)
> > >+		return;
> > >+
> > >+	CS_UNLOCK(drvdata->base);
> > >+
> > >+	debug_os_unlock(drvdata);
> > >+
> > >+	drvdata->edpcsr = readl_relaxed(drvdata->base + EDPCSR);
> > >+
> > >+	/*
> > >+	 * As described in ARM DDI 0487A.k, if the processing
> > >+	 * element (PE) is in debug state, or sample-based
> > >+	 * profiling is prohibited, EDPCSR reads as 0xFFFFFFFF;
> > >+	 * EDCIDSR, EDVIDSR and EDPCSR_HI registers also become
> > >+	 * UNKNOWN state. So directly bail out for this case.
> > >+	 */
> > >+	if (drvdata->edpcsr == EDPCSR_PROHIBITED) {
> > >+		CS_LOCK(drvdata->base);
> > >+		return;
> > >+	}
> > >+
> > >+	/*
> > >+	 * A read of the EDPCSR normally has the side-effect of
> > >+	 * indirectly writing to EDCIDSR, EDVIDSR and EDPCSR_HI;
> > >+	 * at this point it's safe to read value from them.
> > >+	 */
> > 
> > See my comment above about the side effects of memory mapped access.
> 
> Yeah.
> 
> > >+	drvdata->edcidsr = readl_relaxed(drvdata->base + EDCIDSR);
> > >+#ifdef CONFIG_64BIT
> > >+	drvdata->edpcsr_hi = readl_relaxed(drvdata->base + EDPCSR_HI);
> > >+#endif
> > 
> > >+
> > >+	if (drvdata->edvidsr_present)
> > >+		drvdata->edvidsr = readl_relaxed(drvdata->base + EDVIDSR);
> > >+
> > >+	CS_LOCK(drvdata->base);
> > >+}
> > >+
> > 
> > >+#ifndef CONFIG_64BIT
> > 
> > I guess this doesn't help for an ARMv8 32bit only core (e.g, Cortex-A32). And
> > unfortunately, there are conflicting definitions for the values for PCSROffset w.r.t
> > ARMv8 and ARMv7.
> > 
> > DBGDEVID1[3:0] For ARMv7 :
> > 
> > 0000 - Sample offset applies based on the instruction state.
> > 0001 - No offset applies.
> > 
> > EDDEVID1[3:0] For ARMv8 :
> > 0000 - EDPCSR not implemented
> > 0010 - EDPCSR implemented without offsets, but do not use in AArch32 state!
> > 
> > So there is no easy way to make sense of the value, unless you know which version
> > of the architecture is in use. Or may be we could co-relate it with the value from
> > DEVID.
> > 
> > i.e, EDPCSR is not implemented do not register this device, see comments on debug_probe().
> > ( And we should also include the following test for 32bit code to see if edpcsr is implemented.
> > See comments on debug_init_arch_data() )
> > 
> > 
> > That way, we could use the following inference from the PCSROffset value :
> > 
> > 0000 - Sample offset applies based on the instruction state (indicated by PCSR[0])
> > 0001 - No offset applies.
> > 0010 - No offset applies, but do not use in AArch32 mode
> 
> Just now I went through ARM ARM and ARMv8 ARM, this makes sense to me.
> Thanks for good pointing for this.
> 
> > >+static bool debug_pc_has_offset(struct debug_drvdata *drvdata)
> > >+{
> > >+	u32 pcsr_offset;
> > >+
> > >+	pcsr_offset = drvdata->eddevid1 & EDDEVID1_PCSR_OFFSET_MASK;
> > >+
> > >+	return (pcsr_offset == EDDEVID1_PCSR_OFFSET_INS_SET);
> > >+}
> > >+
> > >+static unsigned long debug_adjust_pc(struct debug_drvdata *drvdata,
> > >+				     unsigned long pc)
> > >+{
> > >+	unsigned long arm_inst_offset = 0, thumb_inst_offset = 0;
> > >+
> > >+	if (debug_pc_has_offset(drvdata)) {
> > >+		arm_inst_offset = 8;
> > >+		thumb_inst_offset = 4;
> > >+	}
> > >+
> > >+	/* Handle thumb instruction */
> > >+	if (pc & EDPCSR_THUMB) {
> > >+		pc = (pc & EDPCSR_THUMB_INST_MASK) - thumb_inst_offset;
> > >+		return pc;
> > >+	}
> > >+
> > >+	/*
> > >+	 * Handle arm instruction offset, if the arm instruction
> > >+	 * is not 4 byte alignment then it's possible the case
> > >+	 * for implementation defined; keep original value for this
> > >+	 * case and print info for notice.
> > >+	 */
> > >+	if (pc & BIT(1))
> > >+		pr_emerg("Instruction offset is implementation defined\n");
> > 
> > I am struggling to find the any mention about this in the ARM ARM. Please could
> > you point me to it.
> 
> Sure, please see ARM DDI 0406C.b, chapter C11.11.34 "
> DBGPCSR, Program Counter Sampling Register":
> 
> A profiling tool can use the value of the T bit to calculate the
> instruction address as follows:
> 
> When an offset is applied to the sampled address
> - if T is 0 and DBGPCSR[1] is 0, ((DBGPCSR[31:2] << 2) - 8) is the
> address of the sampled ARM instruction
> - if T is 0 and DBGPCSR[1] is 1, the instruction address is
> IMPLEMENTATION DEFINED
> - if T is 1, ((DBGPCSR[31:1] << 1) - 4) is the address of the sampled
> Thumb or ThumbEE instruction.
> 
> When no offset is applied to the sampled address
> -  if T is 0 and DBGPCSR[1] is 0, (DBGPCSR[31:2] << 2) is the address
> of the sampled ARM instruction
> -  if T is 0 and DBGPCSR[1] is 1, the instruction address is
> IMPLEMENTATION DEFINED
> - if T is 1, (DBGPCSR[31:1] << 1) is the address of the sampled Thumb
> or ThumbEE instruction.
> 
> > >+static void debug_init_arch_data(void *info)
> > >+{
> > >+	struct debug_drvdata *drvdata = info;
> > >+	u32 mode;
> > >+
> > >+	CS_UNLOCK(drvdata->base);
> > >+
> > >+	debug_os_unlock(drvdata);
> > >+
> > >+	/* Read device info */
> > >+	drvdata->eddevid  = readl_relaxed(drvdata->base + EDDEVID);
> > >+	drvdata->eddevid1 = readl_relaxed(drvdata->base + EDDEVID1);
> > >+
> > >+	/* Parse implementation feature */
> > >+	mode = drvdata->eddevid & EDDEVID_PCSAMPLE_MODE;
> > >+	if (mode == EDDEVID_IMPL_FULL) {
> > >+		drvdata->edpcsr_present  = true;
> > >+		drvdata->edvidsr_present = true;
> > >+	} else if (mode == EDDEVID_IMPL_EDPCSR_EDCIDSR) {
> > >+		drvdata->edpcsr_present  = true;
> > >+		drvdata->edvidsr_present = false;
> > 
> > As discussed above, we need to consult the DEVID1:PCSROffset for AArch32 to decide
> > if we have the edpcsr implemented on ARMv8.
> 
> Yeah.
> 
> > >+	} else {
> > >+		drvdata->edpcsr_present  = false;
> > >+		drvdata->edvidsr_present = false;
> > >+	}
> > >+
> > >+	CS_LOCK(drvdata->base);
> > >+}
> > >+
> > >+static int debug_probe(struct amba_device *adev, const struct amba_id *id)
> > >+{
> > >+	void __iomem *base;
> > >+	struct device *dev = &adev->dev;
> > >+	struct debug_drvdata *drvdata;
> > >+	struct resource *res = &adev->res;
> > >+	struct device_node *np = adev->dev.of_node;
> > >+	char buf[32];
> > >+	static int debug_count;
> > >+
> > >+	drvdata = devm_kzalloc(dev, sizeof(*drvdata), GFP_KERNEL);
> > >+	if (!drvdata)
> > >+		return -ENOMEM;
> > >+
> > >+	drvdata->cpu = np ? of_coresight_get_cpu(np) : 0;
> > >+	drvdata->dev = &adev->dev;
> > >+
> > >+	dev_set_drvdata(dev, drvdata);
> > >+
> > >+	/* Validity for the resource is already checked by the AMBA core */
> > >+	base = devm_ioremap_resource(dev, res);
> > >+	if (IS_ERR(base))
> > >+		return PTR_ERR(base);
> > >+
> > >+	drvdata->base = base;
> > >+
> > >+	get_online_cpus();
> > >+	per_cpu(debug_drvdata, drvdata->cpu) = drvdata;
> > >+
> > >+	if (smp_call_function_single(drvdata->cpu,
> > >+				debug_init_arch_data, drvdata, 1))
> > >+		dev_err(dev, "Debug arch init failed\n");
> > 
> > If this fails (say the CPU was offline), should we still return success ?
> > And may be we should check if the drvdata->edpcsr_present to detect if the CPU
> > implements the PC Sampling and return failure here if it doesn't.
> 
> Will fix.
> 
> > >+
> > >+	put_online_cpus();
> > >+
> > >+	if (!debug_count++)
> > >+		atomic_notifier_chain_register(&panic_notifier_list,
> > >+					       &debug_notifier);
> > >+
> > 
> > >+	sprintf(buf, (char *)id->data, drvdata->cpu);
> > >+	dev_info(dev, "%s initialized\n", buf);
> > 
> > This could simply be :
> > 	dev_info(dev, "Coresight debug-CPU%d initialized\n", drvdata->cpu);
> > 
> > and get rid of the static string and the buffer, see below.
> > 
> > >+	return 0;
> > >+}
> > >+
> > >+static struct amba_id debug_ids[] = {
> > >+	{       /* Debug for Cortex-A53 */
> > >+		.id	= 0x000bbd03,
> > >+		.mask	= 0x000fffff,
> > 
> > ...
> > 
> > >+		.data   = "Coresight debug-CPU%d",
> > 
> > I think this is pointless, as the debug area we are interested in is always associated
> > with a CPU, we could as well figure out what to print from the drvdata->cpu above.
> 
> I prefer to follow your suggestion for upper two comments; but I'd like
> check with Mathieu, due I followed up Mathieu's suggestion to write
> current code.

The end result is the same - I'm good either way.

Thanks,
Mathieu

> 
> Thanks,
> Leo Yan

[toc] | [prev] | [next] | [standalone]


#1592219 — [PATCH v3 1/5] coresight: bindings for debug module

FromLeo Yan <leo.yan@linaro.org>
Date2017-03-03 19:50 +0100
Subject[PATCH v3 1/5] coresight: bindings for debug module
Message-ID<th0r9-5c2-35@gated-at.bofh.it>
In reply to#1591705
According to ARMv8 architecture reference manual (ARM DDI 0487A.k)
Chapter 'Part H: External debug', the CPU can integrate debug module
and it can support self-hosted debug and external debug. Especially
for supporting self-hosted debug, this means the program can access
the debug module from mmio region; and usually the mmio region is
integrated with coresight.

So add document for binding debug component, includes binding to APB
clock; and also need specify the CPU node which the debug module is
dedicated to specific CPU.

Suggested-by: Mike Leach <mike.leach@linaro.org>
Reviewed-by: Mathieu Poirier <mathieu.poirier@linaro.org>
Signed-off-by: Leo Yan <leo.yan@linaro.org>
---
 .../devicetree/bindings/arm/coresight-debug.txt    | 40 ++++++++++++++++++++++
 1 file changed, 40 insertions(+)
 create mode 100644 Documentation/devicetree/bindings/arm/coresight-debug.txt

diff --git a/Documentation/devicetree/bindings/arm/coresight-debug.txt b/Documentation/devicetree/bindings/arm/coresight-debug.txt
new file mode 100644
index 0000000..92e5003
--- /dev/null
+++ b/Documentation/devicetree/bindings/arm/coresight-debug.txt
@@ -0,0 +1,40 @@
+* CoreSight Debug Component:
+
+CoreSight debug component are compliant with the ARMv8 architecture reference
+manual (ARM DDI 0487A.k) Chapter 'Part H: External debug'. The external debug
+module is mainly used for two modes: self-hosted debug and external debug, and
+it can be accessed from mmio region from Coresight and eventually the debug
+module connects with CPU for debugging. And the debug module provides
+sample-based profiling extension, which can be used to sample CPU program
+counter, secure state and exception level, etc; usually every CPU has one
+dedicated debug module to be connected.
+
+Required properties:
+
+- compatible : should be
+	     * "arm,coresight-debug"; supplemented with "arm,primecell" since
+	       this driver is using the AMBA bus interface.
+
+- reg : physical base address and length of the register set.
+
+- clocks : the clock associated to this component.
+
+- clock-names : the name of the clock referenced by the code. Since we are
+                using the AMBA framework, the name of the clock providing
+		the interconnect should be "apb_pclk" and the clock is
+		mandatory. The interface between the debug logic and the
+		processor core is clocked by the internal CPU clock, so it
+		is enabled with CPU clock by default.
+
+- cpu : the cpu phandle the debug module is affined to. When omitted
+	the module is considered to belong to CPU0.
+
+Example:
+
+	debug@f6590000 {
+		compatible = "arm,coresight-debug","arm,primecell";
+		reg = <0 0xf6590000 0 0x1000>;
+		clocks = <&sys_ctrl HI6220_DAPB_CLK>;
+		clock-names = "apb_pclk";
+		cpu = <&cpu0>;
+	};
-- 
2.7.4

[toc] | [prev] | [next] | [standalone]


#1596068 — Re: [v3 1/5] coresight: bindings for debug module

FromSuzuki K Poulose <suzuki.poulose@arm.com>
Date2017-03-09 14:30 +0100
SubjectRe: [v3 1/5] coresight: bindings for debug module
Message-ID<tj6iL-5Jn-35@gated-at.bofh.it>
In reply to#1592219
To be honest, coresight-debug sounds too generic and could be confusing with lot
of the other components. To be precise, the area is for External Debug to a CPU.
So "arm,coresight-cpu-debug" or even "arm,coresight-cpu-external-debug" sounds
more appropriate to me.

Suzuki

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web