From patchwork Thu Aug 18 10:02:42 2016
Content-Type: text/plain; charset="utf-8"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
X-Patchwork-Submitter: Ard Biesheuvel <ard.biesheuvel@linaro.org>
X-Patchwork-Id: 74146
Delivered-To: patch@linaro.org
Received: by 10.140.29.52 with SMTP id a49csp264045qga;
 Thu, 18 Aug 2016 03:04:47 -0700 (PDT)
X-Received: by 10.98.149.22 with SMTP id p22mr2554556pfd.88.1471514687522;
 Thu, 18 Aug 2016 03:04:47 -0700 (PDT)
Return-Path: <linux-arm-kernel-bounces+patch=linaro.org@lists.infradead.org>
Received: from bombadil.infradead.org (bombadil.infradead.org.
 [2001:1868:205::9]) by mx.google.com with ESMTPS id
 l24si1815583pfa.120.2016.08.18.03.04.47 for <patch@linaro.org>
 (version=TLS1_2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128/128);
 Thu, 18 Aug 2016 03:04:47 -0700 (PDT)
Received-SPF: pass (google.com: best guess record for domain of
 linux-arm-kernel-bounces+patch=linaro.org@lists.infradead.org
 designates 2001:1868:205::9 as permitted sender)
 client-ip=2001:1868:205::9; 
Authentication-Results: mx.google.com;
 dkim=neutral (body hash did not verify) header.i=@linaro.org;
 spf=pass (google.com: best guess record for domain of
 linux-arm-kernel-bounces+patch=linaro.org@lists.infradead.org
 designates 2001:1868:205::9 as permitted sender)
 smtp.mailfrom=linux-arm-kernel-bounces+patch=linaro.org@lists.infradead.org;
 dmarc=fail (p=NONE dis=NONE) header.from=linaro.org
Received: from localhost ([127.0.0.1] helo=bombadil.infradead.org)
 by bombadil.infradead.org with esmtp (Exim 4.85_2 #1 (Red Hat Linux))
 id 1baKAr-0001EZ-RD; Thu, 18 Aug 2016 10:03:45 +0000
Received: from mail-wm0-x230.google.com ([2a00:1450:400c:c09::230])
 by bombadil.infradead.org with esmtps (Exim 4.85_2 #1 (Red Hat
 Linux)) id 1baKAL-000135-Lq for linux-arm-kernel@lists.infradead.org;
 Thu, 18 Aug 2016 10:03:15 +0000
Received: by mail-wm0-x230.google.com with SMTP id o80so24283991wme.1
 for <linux-arm-kernel@lists.infradead.org>;
 Thu, 18 Aug 2016 03:02:57 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linaro.org; s=google; 
 h=from:to:cc:subject:date:message-id:in-reply-to:references;
 bh=MAGzJlPCariSOh0uDoF0xcOKJZfYfZQ/QtTPg98i2eQ=;
 b=MdevSgoPltdBL5qS9LFV6/mpiipMN5tazaxTFOTMDOTwKCj1k218xURSGx8kRy5D0K
 hxy/wQQWDqyDlwox5l0jOxkS+rr6ZSY8jUxsXxDJaOhQ8ppTGMwkwVhtQMbrsnogKUJM
 8o2CEzqDTGClXdFHwio5TVFKTQERYVyifhvS0=
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
 d=1e100.net; s=20130820;
 h=x-gm-message-state:from:to:cc:subject:date:message-id:in-reply-to
 :references;
 bh=MAGzJlPCariSOh0uDoF0xcOKJZfYfZQ/QtTPg98i2eQ=;
 b=YMXtHniVD1vWojI/S63REaSsCilhXDn+jcT/PMT5g7b2vIQpW4eU8qpi9GFBIKYf/f
 q2XK1hd9b0YAim4389CDVbFhl0K/nWyWj6YTnA0VEhve3RU5jildLZ/Lt5nzZ9lJqtUp
 pzwhhKbwXEo68X/JwV1MnwObCNKTsYrAUyIHR1SEGIj4orJ2fspaeVNi5P3MBZlzDmDX
 ap8dWUxiCt28v5xBlcxJHe1B1GDxx/a9zS6NLECuG2bKX+pPAmBx2LlxD4w74SzaRCXg
 vyjPGmzjxjJI2aSy0FH1OcOlaMrt1k7SCOkeZXhhgpmKJnT/xByMoClCfJBMlN2EJDi/
 y4Eg==
X-Gm-Message-State: AEkoouv/NDdNdXnX2hT8tQD2F6hoxw+3eDO7/fwfXE+nPF9Y8dw3EFB/p72JVdUlMr2M8Rb5
X-Received: by 10.28.212.211 with SMTP id l202mr1814952wmg.109.1471514576363; 
 Thu, 18 Aug 2016 03:02:56 -0700 (PDT)
Received: from localhost.localdomain
 (46.red-81-37-107.dynamicip.rima-tde.net. [81.37.107.46])
 by smtp.gmail.com with ESMTPSA id
 b130sm1860450wmg.19.2016.08.18.03.02.54
 (version=TLS1_2 cipher=ECDHE-RSA-AES128-SHA bits=128/128);
 Thu, 18 Aug 2016 03:02:55 -0700 (PDT)
From: Ard Biesheuvel <ard.biesheuvel@linaro.org>
To: linux-arm-kernel@lists.infradead.org, linux@arm.linux.org.uk,
 dave.martin@arm.com
Subject: [PATCH v3 3/4] ARM: kernel: sort relocation sections before
 allocating PLTs
Date: Thu, 18 Aug 2016 12:02:42 +0200
Message-Id: <1471514563-28172-4-git-send-email-ard.biesheuvel@linaro.org>
X-Mailer: git-send-email 2.7.4
In-Reply-To: <1471514563-28172-1-git-send-email-ard.biesheuvel@linaro.org>
References: <1471514563-28172-1-git-send-email-ard.biesheuvel@linaro.org>
X-CRM114-Version: 20100106-BlameMichelson ( TRE 0.8.0 (BSD) ) MR-646709E3 
X-CRM114-CacheID: sfid-20160818_030314_126135_CDFA6605 
X-CRM114-Status: GOOD (  22.34  )
X-Spam-Score: -2.7 (--)
X-Spam-Report: SpamAssassin version 3.4.0 on bombadil.infradead.org summary:
 Content analysis details:   (-2.7 points)
 pts rule name              description
 ---- ----------------------
 --------------------------------------------------
 -0.7 RCVD_IN_DNSWL_LOW RBL: Sender listed at http://www.dnswl.org/,
 low
 trust [2a00:1450:400c:c09:0:0:0:230 listed in] [list.dnswl.org]
 -0.0 SPF_PASS               SPF: sender matches SPF record
 -1.9 BAYES_00               BODY: Bayes spam probability is 0 to 1%
 [score: 0.0000]
 -0.1 DKIM_VALID Message has at least one valid DKIM or DK signature
 -0.1 DKIM_VALID_AU Message has a valid DKIM or DK signature from
 author's domain
 0.1 DKIM_SIGNED            Message has a DKIM or DK signature,
 not necessarily valid
X-BeenThere: linux-arm-kernel@lists.infradead.org
X-Mailman-Version: 2.1.20
Precedence: list
List-Id: <linux-arm-kernel.lists.infradead.org>
List-Unsubscribe: <http://lists.infradead.org/mailman/options/linux-arm-kernel>, 
 <mailto:linux-arm-kernel-request@lists.infradead.org?subject=unsubscribe>
List-Archive: <http://lists.infradead.org/pipermail/linux-arm-kernel/>
List-Post: <mailto:linux-arm-kernel@lists.infradead.org>
List-Help: <mailto:linux-arm-kernel-request@lists.infradead.org?subject=help>
List-Subscribe: <http://lists.infradead.org/mailman/listinfo/linux-arm-kernel>, 
 <mailto:linux-arm-kernel-request@lists.infradead.org?subject=subscribe>
Cc: namhyung.kim@lge.com, youngho.shin@lge.com, arnd@arndb.de,
 Ard Biesheuvel <ard.biesheuvel@linaro.org>, neidhard.kim@lge.com,
 chanho.min@lge.com
MIME-Version: 1.0
Sender: "linux-arm-kernel" <linux-arm-kernel-bounces@lists.infradead.org>
Errors-To: linux-arm-kernel-bounces+patch=linaro.org@lists.infradead.org

The PLT allocation routines try to establish an upper bound on the
number of PLT entries that will be required at relocation time, and
optimize this by disregarding duplicates (i.e., PLT entries that will
end up pointing to the same function). This is currently a O(n^2)
algorithm, but we can greatly simplify this by
- sorting the relocation section so that relocations that can use the
  same PLT entry will be listed adjacently,
- disregard jump/call relocations with addends; these are highly unusual,
  for relocations against SHN_UNDEF symbols, and so we can simply allocate
  a PLT entry for each one we encounter, without trying to optimize away
  duplicates.

Signed-off-by: Ard Biesheuvel <ard.biesheuvel@linaro.org>
---
 arch/arm/kernel/module-plts.c | 93 ++++++++++++++------
 1 file changed, 64 insertions(+), 29 deletions(-)

-- 
2.7.4


_______________________________________________
linux-arm-kernel mailing list
linux-arm-kernel@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-arm-kernel

diff --git a/arch/arm/kernel/module-plts.c b/arch/arm/kernel/module-plts.c
index 59255d8c9510..5b000ea64fa4 100644
--- a/arch/arm/kernel/module-plts.c
+++ b/arch/arm/kernel/module-plts.c
@@ -9,6 +9,7 @@
 #include <linux/elf.h>
 #include <linux/kernel.h>
 #include <linux/module.h>
+#include <linux/sort.h>
 
 #include <asm/cache.h>
 #include <asm/opcodes.h>
@@ -63,28 +64,54 @@ u32 get_module_plt(struct module *mod, unsigned long loc, Elf32_Addr val)
 	BUG();
 }
 
-static int duplicate_rel(Elf32_Addr base, const Elf32_Rel *rel, int num,
-			   u32 mask)
+#define cmp_3way(a,b)	((a) < (b) ? -1 : (a) > (b))
+
+static int cmp_rel(const void *a, const void *b)
 {
-	u32 *loc1, *loc2;
+	const Elf32_Rel *x = a, *y = b;
 	int i;
 
-	for (i = 0; i < num; i++) {
-		if (rel[i].r_info != rel[num].r_info)
-			continue;
+	/* sort by type and symbol index */
+	i = cmp_3way(ELF32_R_TYPE(x->r_info), ELF32_R_TYPE(y->r_info));
+	if (i == 0)
+		i = cmp_3way(ELF32_R_SYM(x->r_info), ELF32_R_SYM(y->r_info));
+	return i;
+}
 
-		/*
-		 * Identical relocation types against identical symbols can
-		 * still result in different PLT entries if the addend in the
-		 * place is different. So resolve the target of the relocation
-		 * to compare the values.
-		 */
-		loc1 = (u32 *)(base + rel[i].r_offset);
-		loc2 = (u32 *)(base + rel[num].r_offset);
-		if (((*loc1 ^ *loc2) & mask) == 0)
-			return 1;
+static bool is_zero_addend_relocation(u32 *tval)
+{
+	/*
+	 * Do a bitwise compare on the raw addend rather than fully decoding
+	 * the offset and doing an arithmetic comparison.
+	 * Note that a zero-addend jump/call relocation is encoded taking the
+	 * PC bias into account, i.e., -8 for ARM and -4 for Thumb2.
+	 */
+	if (IS_ENABLED(CONFIG_THUMB2_KERNEL)) {
+		u16 upper, lower;
+
+		upper = __mem_to_opcode_thumb16(((u16 *)tval)[0]);
+		lower = __mem_to_opcode_thumb16(((u16 *)tval)[1]);
+
+		return (upper & 0x7ff) == 0x7ff && (lower & 0x2fff) == 0x2ffe;
 	}
-	return 0;
+	return (__mem_to_opcode_arm(*tval) & 0xffffff) == 0xfffffe;
+}
+
+static bool duplicate_rel(Elf32_Addr base, const Elf32_Rel *rel, int num)
+{
+	const Elf32_Rel *prev;
+
+	/*
+	 * Entries are sorted by type and symbol index. That means that,
+	 * if a duplicate entry exists, it must be in the preceding
+	 * slot.
+	 */
+	if (!num)
+		return false;
+
+	prev = rel + num - 1;
+	return cmp_rel(rel + num, prev) == 0 &&
+	       is_zero_addend_relocation((u32 *)(base + prev->r_offset));
 }
 
 /* Count how many PLT entries we may need */
@@ -93,20 +120,12 @@ static unsigned int count_plts(const Elf32_Sym *syms, Elf32_Addr base,
 {
 	unsigned int ret = 0;
 	const Elf32_Sym *s;
-	u32 mask;
 	int i;
 
-	if (IS_ENABLED(CONFIG_THUMB2_KERNEL))
-		mask = __opcode_to_mem_thumb32(0x07ff2fff);
-	else
-		mask = __opcode_to_mem_arm(0x00ffffff);
-
-	/*
-	 * Sure, this is order(n^2), but it's usually short, and not
-	 * time critical
-	 */
 	for (i = 0; i < num; i++) {
 		switch (ELF32_R_TYPE(rel[i].r_info)) {
+			u32 *tval;
+
 		case R_ARM_THM_CALL:
 		case R_ARM_THM_JUMP24:
 			BUG_ON(!IS_ENABLED(CONFIG_THUMB2_KERNEL));
@@ -125,7 +144,20 @@ static unsigned int count_plts(const Elf32_Sym *syms, Elf32_Addr base,
 			if (s->st_shndx != SHN_UNDEF)
 				break;
 
-			if (!duplicate_rel(base, rel, i, mask))
+			/*
+			 * Jump relocations with non-zero addends against
+			 * undefined symbols are supported by the ELF spec, but
+			 * do not occur in practice (e.g., 'jump n bytes past
+			 * the entry point of undefined function symbol f').
+			 * So we need to support them, but there is no need to
+			 * take them into consideration when trying to optimize
+			 * this code. So let's only check for duplicates when
+			 * the addend is zero.
+			 */
+			tval = (u32 *)(base + rel[i].r_offset);
+
+			if (!is_zero_addend_relocation(tval) ||
+			    !duplicate_rel(base, rel, i))
 				ret++;
 		}
 	}
@@ -160,7 +192,7 @@ int module_frob_arch_sections(Elf_Ehdr *ehdr, Elf_Shdr *sechdrs,
 	}
 
 	for (s = sechdrs + 1; s < sechdrs_end; ++s) {
-		const Elf32_Rel *rels = (void *)ehdr + s->sh_offset;
+		Elf32_Rel *rels = (void *)ehdr + s->sh_offset;
 		int numrels = s->sh_size / sizeof(Elf32_Rel);
 		Elf32_Shdr *dstsec = sechdrs + s->sh_info;
 
@@ -171,6 +203,9 @@ int module_frob_arch_sections(Elf_Ehdr *ehdr, Elf_Shdr *sechdrs,
 		if (!(dstsec->sh_flags & SHF_EXECINSTR))
 			continue;
 
+		/* sort by type and symbol index */
+		sort(rels, numrels, sizeof(Elf32_Rel), cmp_rel, NULL);
+
 		plts += count_plts(syms, dstsec->sh_addr, rels, numrels);
 	}