00:23 zmike: https://i.imgflip.com/5pojqq.png
05:57 Lynne: I wish I can send validation layer drive-by patches by shouting diffs at someone or something
06:24 tzimmermann: airlied, sima, hi. please backmerge 7.3-rc4 into drm-next for later backporting into drm-misc-next
06:26 sima: oops sorry forgot yesterday
06:27 sima: tzimmermann, did phasta tell you the commit he needs?
06:27 tzimmermann: no
06:27 tzimmermann: sry
06:32 sima: hm wait until he's back on irc so I can type it up?
06:34 sima: I'm doing the conflict resolving meanwhile
06:35 sima: and breakfast
06:35 tzimmermann: sure
06:36 tzimmermann: sima, i just noticed, phasta actually sent me the commit: "If we'd need 2ab510e63197360945f915dd5631a77c63ac6b27 from drm-misc-fixes in drm-misc-next, what would the process be like?"
06:39 airlied: I'm just merging the last xe one before I did it
06:40 sima: airlied, ok I'll leave it to you then
06:40sima:back to breakfast
06:46 HdkR: 3
07:46 dolphin: airlied, sima: looks like the FUSE discussion stalled and currently the popular opinion seems to be that DRM subsystem should accommodate unbounded page-fault times
07:51 sima: no
07:52 sima: like not sure what to add, but I don't think there's any reasonable engineering where this makes sense
07:52 sima: especially since it's documented in the fuse docs that fuse might randomly hang/stall/whatever
07:54 sima: like I think best we can reasonably do is what the net fs do when the network disappears and make sure that all these things are killable
07:54 sima: but that would be up to fuse to make sure that fuse page faults are killable in all cases where they might be stuck
07:55 sima: once the forever-stuck page faults are all killable we should be able to at least get out of any bind through the (hopefully verified through nasty igt tests) error paths we need anyway to handle faults for any user access
07:58 sima: like no as in https://imgflip.com/i/b1q0ng
08:00 sima: dolphin, so I'd just use this to make triple sure igt can cover even the worst cases (like the fault happening as part of a user fence completion work) and then call it done
08:01 sima: we need to ensure that anyway, and I think that's enough to formally handle the "fuse can get stuck" requirement
08:02 dolphin: right, we should absolutely be ready to handle errors during copy_from_user and userptr
08:03 dolphin: however, as far as I can tell, killability question becomes, copy_from_user() may never actually return before user goes and kills the bad FUSE connection
08:03 sima: the one caveat is that from kernel workers the kill signal obviously wont work to get the thread unstuck, but fuse has the knob to nuke an fs
08:03 sima: yeah
08:03 sima: but I think net fs have the same problem if you get stuck in async workers and things like that, but not sure
08:05 sima: also even if we (= overall kernel community) sort this out with a timeout I feel like this should be a per-mm setting
08:05 sima: the app knows best what it can take I think
08:05 sima: with maybe a per-subsystem max or something like that
08:06 dolphin: yeah, I'm heavily agreeing, it just sounds like others are not :)
08:06 sima: but given that copy_*_user_timeout is not a thing it's very clear that this isn't how it's supposed to work
08:06 dolphin: FUSE also refusing to put in default timeout because can't know the maximum, so I guess we just have a kernel that locks up :P
08:07 sima: dolphin, imo the complete lack of any _timeout variant for user access makes this a crystal clear case
08:07 sima: callers _cannot_ be responsible for the timeout, at least under the current uaccess.h contract
08:07 sima: case closed
08:08 dolphin: I guess they're not arguin caller to imply timeout, just that refactor everything so that only single process will lock up
08:08 dolphin: same process as the owner of the FUSE
08:09 dolphin: which I don't think will be too good either
08:10 sima: but that means we can't hold any shared locks
08:10 sima: well any locks, since generally they exist because something is shared
08:10 dolphin: yeah, only per-process locks
08:11 sima: that seems reasonable if you write an fs or storage driver, since that's the rules already mostly anyway since a page fault can recurse
08:11 sima: but for anything else it's just way too strict and not how things have been done
08:12 dolphin: right, I'm pretty sure i915 predates fuse, for example :)
08:12 sima: oh, any of the driver subsystems that don't have storage or networking essentially
08:12 dolphin: and the early HW implementations, surely were not per-process isolating
08:16 sima: hm, _killable for net fs is almost 2 decades old, so maybe you can make a case that this issue predates i915
08:16 sima: but eh
08:18 sima: essentially it would mean that any user access from worker threads would need the _inatomic variant, and if there's a potential fault you need to start a one-off work item that does not hold any locks or block anything to proceed for writes
08:18 sima: and for reads drop all locks except mm/process-local ones I guess?
08:19 sima: but this all falls apart because guc has a single ring buffer for communication and page faults are ordered/blocking (I think so at least still)
08:19 sima: so we must be able to get at mm locks for forward progress without dying, and that's just not possible under this constraint
08:20 sima: hence back to "not our problem" I think
08:25 sima: agd5f, maybe need amd's take more, the thread seems to have died out a bit https://lore.kernel.org/dri-devel/178896547895.102087.5758513424678942429@jlahtine-mobl/
08:32 airlied: phasta: what reason did you want a backmerge from rc4 to -next?
08:33 phasta: airlied: 2ab510e63197360945f915dd5631a77c63ac6b27 would be needed as a groundlayer for further patches that extend the locking in drm_sched
08:33 phasta: (from misc-fixes)
08:35 airlied: phasta: great, thanks
08:35 phasta: airlied: thank you! Hope there's no trouble
08:38 airlied: just pushing it to drm-next now so seems fine
08:38 airlied: tzimmermann: pushed now
10:09 dolphin: sima: you should convert the above image to ascii and respond to the m-l? :)
10:09 sima: ...
10:09 sima: not sure that'll help :-D
10:10 sima: slightly more seriously I've pondered how this would all even work when trying to do user access or page faults from kernel threads
10:10 sima: since you cannot interrupt them with signals
10:11 sima: like we'd need _timeout variants of all these (and really all, so uaccess, gup, hmm, ...) and maybe a timeout like the gpu preempt timeout
10:11 sima: which just feels like bad bandaid
10:11 sima: and I didn't come up with anything better
10:12 sima: like maybe some horrible kludge that kill signals to a process also get delivered to attached kernel workers that used kthread_use_mm, but that's so much infra
12:10 tzimmermann: thanks, airlied
13:03 tzimmermann: phasta, drm-misc-next now has commit 2ab510e63197
13:14 dolphin: sima: I'll definitely appreciate to pull in more folks to the discussion, I can of course go and reiterate my stance but I have doubts it'll make a difference
14:16 phasta: tzimmermann: very cool, thx guys
15:06 karolherbst: valpackett: mind giving this a try for the new libclc-23 locations? https://gitlab.freedesktop.org/mesa/mesa/-/merge_requests/44624
17:03 cwabbott: apparently mesa CI doesn't correctly check whether compilation succeeds sometimes: https://gitlab.freedesktop.org/mesa/mesa/-/jobs/110866566