Migrate Layer

The Xe migrate layer is used generate jobs which can copy memory (eviction), clear memory, or program tables (binds). This layer exists in every GT, has a migrate engine, and uses a special VM for all generated jobs.

Special VM details

The special VM is configured with a page structure where we can dynamically map BOs which need to be copied and cleared, dynamically map other VM’s page table BOs for updates, and identity map the entire device’s VRAM with 1 GB pages.

Currently the page structure consists of 32 physical pages with 16 being reserved for BO mapping during copies and clear, 1 reserved for kernel binds, several pages are needed to setup the identity mappings (exact number based on how many bits of address space the device has), and the rest are reserved user bind operations.

TODO: Diagram of layout

Bind jobs

A bind job consist of two batches and runs either on the migrate engine (kernel binds) or the bind engine passed in (user binds). In both cases the VM of the engine is the migrate VM.

The first batch is used to update the migration VM page structure to point to the bind VM page table BOs which need to be updated. A physical page is required for this. If it is a user bind, the page is allocated from pool of pages reserved user bind operations with drm_suballoc managing this pool. If it is a kernel bind, the page reserved for kernel binds is used.

The first batch is only required for devices without VRAM as when the device has VRAM the bind VM page table BOs are in VRAM and the identity mapping can be used.

The second batch is used to program page table updated in the bind VM. Why not just one batch? Well the TLBs need to be invalidated between these two batches and that only can be done from the ring.

When the bind job complete, the page allocated is returned the pool of pages reserved for user bind operations if a user bind. No need do this for kernel binds as the reserved kernel page is serially used by each job.

Copy / clear jobs

A copy or clear job consist of two batches and runs on the migrate engine.

Like binds, the first batch is used update the migration VM page structure. In copy jobs, we need to map the source and destination of the BO into page the structure. In clear jobs, we just need to add 1 mapping of BO into the page structure. We use the 16 reserved pages in migration VM for mappings, this gives us a maximum copy size of 16 MB and maximum clear size of 32 MB.

The second batch is used do either do the copy or clear. Again similar to binds, two batches are required as the TLBs need to be invalidated from the ring between the batches.

More than one job will be generated if the BO is larger than maximum copy / clear size.

Future work

Update copy and clear code to use identity mapped VRAM.

Can we rework the use of the pages async binds to use all the entries in each page?

Using large pages for sysmem mappings.

Is it possible to identity map the sysmem? We should explore this.

Migration jobs issued on behalf of GPU page faults and SVM prefetches sit directly in the critical path of a stalled GPU workload. The dominant cost of such a job is not the copy or clear itself but the submission latency: the H2G round trip to GuC, the GuC scheduling decision, and the hardware context switch required to place the migration LRC on an engine.

ULLS removes that cost by keeping the migration context resident and running on the hardware engine across jobs. Instead of the ring going empty and the context being switched out between jobs, the tail of every ULLS job parks the engine on a semaphore wait for the next job’s semaphore, and then advances the ring tail itself. Submitting the next job therefore costs the CPU a single write to signal that semaphore - no H2G, no GuC round trip, no context switch, no MMIO.

ULLS is only used on dGFX platforms with USM support, where a hardware engine is reserved exclusively for migration jobs. Because the engine spins on a semaphore while ULLS is active, it cannot be shared with user submissions.

ULLS can be disabled by setting the migrate_ulls_period_ms configfs attribute to 0. Otherwise, the same attribute controls how long ULLS remains active before exiting, in milliseconds.

A job updates the ring tail to cover its successor, but it is emitted long before that successor exists, so it can not know how much ring the successor will occupy. Every ULLS job is therefore padded out to exactly ULLS_JOB_SIZE_BYTES, which lets the next tail be computed arithmetically from where the current job started.

This is why the shorter jobs still have to reach the same size: the “last” job skips the batch buffers and the postamble, and pads the difference with MI_NOOP. The “first” job is not covered by any predecessor’s tail update and so is unconstrained, but is padded anyway to keep the arithmetic uniform.

Leaving ULLS mode always goes through a “last” job, which emits no tail update, so an ordinary variable length migration job never follows a prediction.

The semaphores live in the driver-defined portion of the migration LRC’s PPHWSP (see LRC_ULLS_PPHWSP_OFFSET, mutually exclusive with the parallel submission area). There are LRC_MIGRATION_ULLS_SEMAPHORE_COUNT of them and a job’s semaphore is selected by seqno % COUNT, so the semaphore ring wraps with the job seqnos. To guarantee a job can never overwrite the semaphore of a job still in flight, the GuC backend caps the migration queue’s scheduler job count at LRC_MIGRATION_ULLS_SEMAPHORE_COUNT - 1.

Emitted by emit_migration_job_gen12() in xe_ring_ops.c:

preamble:       clear semaphore[seqno]  (reuse for a later wrap)
<copy timestamp, start seqno store>
<batch buffer start(s)>                 (skipped on first/last job)
<seqno write + user interrupt>
postamble:      SDI saved ring tail = end of next job
                LRI RING_TAIL = end of next job
                wait on semaphore[seqno + 1]
                                        (skipped on the last job)
pad:            MI_NOOP up to ULLS_JOB_SIZE_DW

The preamble clears the current job’s semaphore so it can be reused once the seqno space wraps. The postamble is what keeps the engine busy: it advances the ring tail over the next job and then blocks on that job’s semaphore, which is only signaled when the job is actually submitted. It advances the saved tail as well as the tail register, keeping the two in step without any help from the CPU, so a context save and restore can not rewind the tail behind work which has already been published.

The tail register write must be non-posted, i.e. it must not carry MI_LRI_FORCE_POSTED. Posted, the new tail is free to land after the command streamer has already drained the rest of the job, at which point the command streamer sees head == the old tail and parks as though the ring were empty. A parked context can be switched off the hardware, and the fast path below has no H2G with which to ask GuC to bring it back.

The tail is published ahead of the semaphore wait rather than after it so that the non-posted write drains while the engine is parked anyway, keeping a register round trip off the path between the semaphore being signaled and the next job running.

In submit_exec_queue() (xe_guc_submit.c), a ULLS job that is not the first one reduces to:

xe_lrc_set_ulls_semaphore(lrc, seqno);          release previous job

The XE_GUC_ACTION_SCHED_CONTEXT H2G is suppressed, and so is the write of the saved ring tail: the previous job’s postamble has already published this job’s tail both in the tail register and in the context image, so the semaphore signal is all that is left. The previous job’s semaphore wait is satisfied and the engine walks straight into this job.

This does assume the context stays resident for as long as ULLS mode is active. Nothing else is scheduled on the reserved engine, so the only ways off the hardware are the “last” job below, or a reset - and a migration job failing already wedges the device.

xe_migrate_ulls_enter() is called from the page fault handler and from the SVM prefetch path, i.e. exactly where low latency migration matters. It takes a PM runtime reference (the device must not suspend while the engine spins), then submits a “first” ULLS job. That first job carries no batch buffer; it exists only to get the context onto the hardware through the normal GuC path and to leave the engine waiting on the next semaphore, pipelining the GuC/HW context switch out of the critical path.

No forcewake reference is required. Nothing in the fast path touches MMIO, and the engine keeps itself awake for as long as it is executing the ring. Not needing host MMIO access is also what lets ULLS run on SRIOV VFs.

Keeping an engine spinning costs power, so ULLS is not left enabled indefinitely. Every enter and every ULLS job submission re-arms xe_migrate.ulls.exit_work with a ULLS_EXIT_JIFFIES delay. When it fires with the queue idle, it submits a “last” ULLS job - again with no batch buffer and, crucially, with no postamble semaphore wait or tail update - which lets the ring drain so the context can be switched off the hardware. The PM reference is then dropped. If the queue was not idle, the worker simply re-arms itself.

The state above is communicated to the ring ops and GuC backend via xe_sched_job.ulls, set under xe_migrate.job_mutex:

  • ULLS_NONE: job submitted outside of ULLS mode

  • ULLS_ENTER: job that enters ULLS mode

  • ULLS_ACTIVE: job submitted while in ULLS mode

  • ULLS_EXIT: job that exits ULLS mode