Improve slow path performance for allocation (#143)

* Remote dealloc refactor.

* Improve remote dealloc

Change remote to count down to 0, so fast path does not need a constant.

Use signed value so that branch does not depend on addition.

* Inline remote_dealloc

The fast path of remote_dealloc is sufficiently compact that it can be
inlined.

* Improve fast path in Slab::alloc

Turn the internal structure into tail calls, to improve fast path.
Should be no algorithmic changes.

* Refactor initialisation to help fast path.

Break lazy initialisation into two functions, so it is easier to codegen
fast paths.

* Minor tidy to statically sized dealloc.

* Refactor semi-slow path for alloc

Make the backup path a bit faster.  Only algorithmic change is to delay
checking for first allocation. Otherwise, should be unchanged.

* Test initial operation of a thread

The first operation a new thread takes is special.  It results in
allocating an allocator, and swinging it into the TLS.  This makes
this a very special path, that is rarely tested.  This test generates
a lot of threads to cover the first alloc and dealloc operations.

* Correctly handle reusing get_noncachable

* Fix large alloc stats

Large alloc stats aren't necessarily balanced on a thread, this changes
to tracking individual pushs and pops, rather than the net effect
(with an unsigned value).

* Fix TLS init on large alloc path

* Add Bump ptrs to allocator

Each allocator has a bump ptr for each size class.  This is no longer
slab local.

Slabs that haven't been fully allocated no longer need to be in the DLL
for this sizeclass.

* Change to a cycle non-empty list

This change reduces the branching in the case of finding a new free
list. Using a non-empty cyclic list enables branch free add, and a
single branch in remove to detect the empty case.

* Update differences

* Rename first allocation

Use needs initialisation as makes more sense for other scenarios.

* Use a ptrdiff to help with zero init.

* Make GlobalPlaceholder zero init

The GlobalPlaceholder allocator is now a zero init block of memory.
This removes various issues for when things are initialised. It is made read-only
to we detect write to it on some platforms.
This commit is contained in:
Matthew Parkinson
2020-03-31 09:17:53 +01:00
committed by GitHub
parent ecef894525
commit d900e29424
20 changed files with 690 additions and 239 deletions

View File

@@ -62,7 +62,7 @@ namespace snmalloc
* This struct is used to represent callbacks for notification from the
* platform. It contains a next pointer as client is responsible for
* allocation as we cannot assume an allocator at this point.
**/
*/
struct PalNotificationObject
{
std::atomic<PalNotificationObject*> pal_next;
@@ -72,12 +72,12 @@ namespace snmalloc
/***
* Wrapper for managing notifications for PAL events
**/
*/
class PalNotifier
{
/**
* List of callbacks to notify
**/
*/
std::atomic<PalNotificationObject*> callbacks = nullptr;
public:
@@ -86,7 +86,7 @@ namespace snmalloc
*
* The object should never be deallocated by the client after calling
* this.
**/
*/
void register_notification(PalNotificationObject* callback)
{
callback->pal_next = nullptr;
@@ -105,7 +105,7 @@ namespace snmalloc
/**
* Calls the pal_notify of all the registered objects.
**/
*/
void notify_all()
{
PalNotificationObject* curr = callbacks;

View File

@@ -34,7 +34,7 @@ namespace snmalloc
/**
* List of callbacks for low-memory notification
**/
*/
static inline PalNotifier low_memory_callbacks;
/**
@@ -98,7 +98,7 @@ namespace snmalloc
* Register callback object for low-memory notifications.
* Client is responsible for allocation, and ensuring the object is live
* for the duration of the program.
**/
*/
static void
register_for_low_memory_callback(PalNotificationObject* callback)
{