* Remote dealloc refactor. * Improve remote dealloc Change remote to count down to 0, so fast path does not need a constant. Use signed value so that branch does not depend on addition. * Inline remote_dealloc The fast path of remote_dealloc is sufficiently compact that it can be inlined. * Improve fast path in Slab::alloc Turn the internal structure into tail calls, to improve fast path. Should be no algorithmic changes. * Refactor initialisation to help fast path. Break lazy initialisation into two functions, so it is easier to codegen fast paths. * Minor tidy to statically sized dealloc. * Refactor semi-slow path for alloc Make the backup path a bit faster. Only algorithmic change is to delay checking for first allocation. Otherwise, should be unchanged. * Test initial operation of a thread The first operation a new thread takes is special. It results in allocating an allocator, and swinging it into the TLS. This makes this a very special path, that is rarely tested. This test generates a lot of threads to cover the first alloc and dealloc operations. * Correctly handle reusing get_noncachable * Fix large alloc stats Large alloc stats aren't necessarily balanced on a thread, this changes to tracking individual pushs and pops, rather than the net effect (with an unsigned value). * Fix TLS init on large alloc path * Add Bump ptrs to allocator Each allocator has a bump ptr for each size class. This is no longer slab local. Slabs that haven't been fully allocated no longer need to be in the DLL for this sizeclass. * Change to a cycle non-empty list This change reduces the branching in the case of finding a new free list. Using a non-empty cyclic list enables branch free add, and a single branch in remove to detect the empty case. * Update differences * Rename first allocation Use needs initialisation as makes more sense for other scenarios. * Use a ptrdiff to help with zero init. * Make GlobalPlaceholder zero init The GlobalPlaceholder allocator is now a zero init block of memory. This removes various issues for when things are initialised. It is made read-only to we detect write to it on some platforms.
120 lines
3.1 KiB
C++
120 lines
3.1 KiB
C++
#pragma once
|
|
|
|
#include "../ds/defines.h"
|
|
|
|
#include <atomic>
|
|
|
|
namespace snmalloc
|
|
{
|
|
/**
|
|
* Flags in a bitfield of optional features that a PAL may support. These
|
|
* should be set in the PAL's `pal_features` static constexpr field.
|
|
*/
|
|
enum PalFeatures : uint64_t
|
|
{
|
|
/**
|
|
* This PAL supports low memory notifications. It must implement a
|
|
* `register_for_low_memory_callback` method that allows callbacks to be
|
|
* registered when the platform detects low-memory and an
|
|
* `expensive_low_memory_check()` method that returns a `bool` indicating
|
|
* whether low memory conditions are still in effect.
|
|
*/
|
|
LowMemoryNotification = (1 << 0),
|
|
/**
|
|
* This PAL natively supports allocation with a guaranteed alignment. If
|
|
* this is not supported, then we will over-allocate and round the
|
|
* allocation.
|
|
*
|
|
* A PAL that does supports this must expose a `request()` method that takes
|
|
* a size and alignment. A PAL that does *not* support it must expose a
|
|
* `request()` method that takes only a size.
|
|
*/
|
|
AlignedAllocation = (1 << 1),
|
|
/**
|
|
* This PAL natively supports lazy commit of pages. This means have large
|
|
* allocations and not touching them does not increase memory usage. This is
|
|
* exposed in the Pal.
|
|
*/
|
|
LazyCommit = (1 << 2),
|
|
};
|
|
/**
|
|
* Flag indicating whether requested memory should be zeroed.
|
|
*/
|
|
enum ZeroMem
|
|
{
|
|
/**
|
|
* Memory should not be zeroed, contents are undefined.
|
|
*/
|
|
NoZero,
|
|
/**
|
|
* Memory must be zeroed. This can be lazily allocated via a copy-on-write
|
|
* mechanism as long as any load from the memory returns zero.
|
|
*/
|
|
YesZero
|
|
};
|
|
|
|
/**
|
|
* Default Tag ID for the Apple class
|
|
*/
|
|
static const int PALAnonDefaultID = 241;
|
|
|
|
/**
|
|
* This struct is used to represent callbacks for notification from the
|
|
* platform. It contains a next pointer as client is responsible for
|
|
* allocation as we cannot assume an allocator at this point.
|
|
*/
|
|
struct PalNotificationObject
|
|
{
|
|
std::atomic<PalNotificationObject*> pal_next;
|
|
|
|
void (*pal_notify)(PalNotificationObject* self);
|
|
};
|
|
|
|
/***
|
|
* Wrapper for managing notifications for PAL events
|
|
*/
|
|
class PalNotifier
|
|
{
|
|
/**
|
|
* List of callbacks to notify
|
|
*/
|
|
std::atomic<PalNotificationObject*> callbacks = nullptr;
|
|
|
|
public:
|
|
/**
|
|
* Register a callback object to be notified
|
|
*
|
|
* The object should never be deallocated by the client after calling
|
|
* this.
|
|
*/
|
|
void register_notification(PalNotificationObject* callback)
|
|
{
|
|
callback->pal_next = nullptr;
|
|
|
|
auto prev = &callbacks;
|
|
auto curr = prev->load();
|
|
do
|
|
{
|
|
while (curr != nullptr)
|
|
{
|
|
prev = &(curr->pal_next);
|
|
curr = prev->load();
|
|
}
|
|
} while (!prev->compare_exchange_weak(curr, callback));
|
|
}
|
|
|
|
/**
|
|
* Calls the pal_notify of all the registered objects.
|
|
*/
|
|
void notify_all()
|
|
{
|
|
PalNotificationObject* curr = callbacks;
|
|
while (curr != nullptr)
|
|
{
|
|
curr->pal_notify(curr);
|
|
curr = curr->pal_next;
|
|
}
|
|
}
|
|
};
|
|
} // namespace snmalloc
|