yslalan Posted Thursday at 09:48 PM Posted Thursday at 09:48 PM https://github.com/LevinAi-arch/rtx-5070ti-laptop-160w-power-limit/releases/tag/v1.8.0 Interesting repo. Just modified a few lines of code to make it work with the RTX PRO. Dell Pro Max 16 Plus Ultra 9-285HX - NVIDIA RTX Pro 4000 Blackwell - 64G DDR5 - UHD+ Display - 3840*2400 OLED - 7T NVMe Dell XPS 16 DA16260 Ultra X7-358H - Arc B390 - 32G LPDDR5X - 3.2K OLED - 1T NVMe
SuperMG Posted Thursday at 11:19 PM Posted Thursday at 11:19 PM 1 hour ago, yslalan said: https://github.com/LevinAi-arch/rtx-5070ti-laptop-160w-power-limit/releases/tag/v1.8.0 Interesting repo. Just modified a few lines of code to make it work with the RTX PRO. Hello. Can this work for RTX Ada Lovelace mobile cards? Because I found this: https://github.com/timmyy123/nvidia-power-control
yslalan Posted Thursday at 11:47 PM Posted Thursday at 11:47 PM Ada and Blackwell should both work. You may just need to add logic to detect the device name properly. 27 minutes ago, SuperMG said: Hello. Can this work for RTX Ada Lovelace mobile cards? Because I found this: https://github.com/timmyy123/nvidia-power-control Dell Pro Max 16 Plus Ultra 9-285HX - NVIDIA RTX Pro 4000 Blackwell - 64G DDR5 - UHD+ Display - 3840*2400 OLED - 7T NVMe Dell XPS 16 DA16260 Ultra X7-358H - Arc B390 - 32G LPDDR5X - 3.2K OLED - 1T NVMe
SuperMG Posted Friday at 12:07 AM Posted Friday at 12:07 AM 20 minutes ago, yslalan said: Ada and Blackwell should both work. You may just need to add logic to detect the device name properly. No way my 4090 150W MXM card can do 175W+ easily on my Clevo? But I would get limited by the power resistors (R006)... And power limiting could work too... Back then we used the bugged Nvidia drivers to power limit our mobile GPUs.
delboi Posted Sunday at 02:16 PM Posted Sunday at 02:16 PM I’m investigating this further with the recurrent WHEA-Logger Event 17 / PCIe AER events on a Dell Pro Max 16 Plus MB16250 with Core Ultra 9 285HX. They occur during early Windows boot across multiple unrelated PCIe endpoints: Intel BE200 Wi-Fi, NVMe SSD and NVIDIA GPU. Raw WHEA/CPER records show the same AER signature: UncorrectableErrorStatus = 0x100000 = Unsupported Request CorrectableErrorStatus = 0xa000 = Advisory Non-Fatal + Header Log Overflow Captured TLP headers decode as PCIe Configuration Read Type 0 requests, apparently probing absent functions/BDFs during early PCI enumeration. Some requests occur only microseconds apart. Linux PCI enumeration source: https://github.com/torvalds/linux/blob/master/drivers/pci/probe.c Linux PCIe AER documentation: https://www.kernel.org/doc/html/latest/PCI/pcieaer-howto.html Intel Arrow Lake Series 2 specification update / errata: https://edc.intel.com/content/www/us/en/design/products/platforms/details/arrow-lake-s/core-ultra-200s-series-processors-specification-update/errata-details/ Relevant Intel erratum: ARL068 — PCIe Gen5 Link Exit from L1 Sub-state Low Power State. It specifically mentions short back-to-back PCIe configuration-space accesses as one trigger condition. Dell also has an official article stating that boot-time WHEA-Logger ID17 is “expected behavior as per Intel” on certain earlier Intel HX systems. Note that Dell’s listed affected systems are 12th-gen/HX platforms, not the MB16250/285HX, so it does not by itself prove this newer system is expected to behave the same way: https://www.dell.com/support/kbdoc/en-uk/000216115/laptops-with-12th-gen-and-12th-gen-hx-intel-core-processors-may-display-warning-message-whea-loggerid17 WHEA-17 itself is quite generic, and I think the associated symptoms may be more significant than the event alone. In my case, I’ve seen intermittent freezing and occasional mouse slowdowns. Dell support has also told me that some users experience BSODs or system crashes. What I’m really looking for here is other people’s ideas, observations and possible avenues to investigate
whyshchuck Posted Monday at 10:56 AM Posted Monday at 10:56 AM 20 hours ago, delboi said: I’m investigating this further with the recurrent WHEA-Logger Event 17 / PCIe AER events on a Dell Pro Max 16 Plus MB16250 with Core Ultra 9 285HX. They occur during early Windows boot across multiple unrelated PCIe endpoints: Intel BE200 Wi-Fi, NVMe SSD and NVIDIA GPU. (...) WHEA-17 itself is quite generic, and I think the associated symptoms may be more significant than the event alone. In my case, I’ve seen intermittent freezing and occasional mouse slowdowns. Dell support has also told me that some users experience BSODs or system crashes. What I’m really looking for here is other people’s ideas, observations and possible avenues to investigate Answering your hardware-swap question - sibling model here (Pro Max 18 Plus, same 285HX CPU): - System board (motherboard): replaced. No change. - Discrete NVIDIA GPU module: physically replaced. No change. - Whole unit: I was given a different unit (fresh Windows, different RAM/disk/ panel). Same behaviour returned - and two units showed an identical WHEA-17 signature side by side. - BE200 Wi-Fi: not physically removed in my case, so I can't speak to that. One difference worth noting: my sustained AER doesn't come from Wi-Fi/NVMe/GPU at boot like yours - it's from the USB4/Thunderbolt host-router downstream port (PCI\VEN_8086&DEV_5786), i.e. my dock link. So for me the BE200 doesn't look central, though the common 285HX / ARL068 angle still fits. Net: board swap, GPU-module swap and a full-unit swap all failed to fix it, which points at the platform rather than a single removable part.
delboi Posted Monday at 11:04 PM Posted Monday at 11:04 PM 12 hours ago, whyshchuck said: Answering your hardware-swap question - sibling model here (Pro Max 18 Plus, same 285HX CPU): - System board (motherboard): replaced. No change. - Discrete NVIDIA GPU module: physically replaced. No change. - Whole unit: I was given a different unit (fresh Windows, different RAM/disk/ panel). Same behaviour returned - and two units showed an identical WHEA-17 signature side by side. - BE200 Wi-Fi: not physically removed in my case, so I can't speak to that. One difference worth noting: my sustained AER doesn't come from Wi-Fi/NVMe/GPU at boot like yours - it's from the USB4/Thunderbolt host-router downstream port (PCI\VEN_8086&DEV_5786), i.e. my dock link. So for me the BE200 doesn't look central, though the common 285HX / ARL068 angle still fits. Net: board swap, GPU-module swap and a full-unit swap all failed to fix it, which points at the platform rather than a single removable part. Thanks for the response I am starting to think this is a fault domain which a PCIe signal integrity issue somewhere between: The remaining fault domains are approximately: CPU PCIe Root Complex / Root Port hardware — including the host-side PCIe PHY/receiver. Motherboard PCIe channel — traces, connectors and other physical elements between the Root Port and endpoint. Power delivery affecting the PCIe subsystem. PCIe reference clock / clocking behaviour. Firmware/software-controlled PCIe behaviour — UEFI/BIOS, ACPI/platform power management, link-state management, link training/retraining and related platform configuration. Your full-unit replacement result is particularly interesting to me because Dell is currently discussing replacing my MB16250 with a Certified Refurbished unit. Was your original machine new or refurbished, and was the complete replacement Dell supplied new or refurbished? If refurbished, do you know anything about its previous repair history? One concern I have is what Dell's refurbishment validation actually tests at PCIe level. I have been using PCIe Lane Margining at the Receiver (LMR) under Linux as part of my investigation. My understanding is that this interface exists specifically to assess receiver timing/voltage margin rather than merely checking that a PCIe device enumerates, negotiates its expected link speed/width and passes ordinary functional tests. I don't know whether Dell performs LMR or equivalent PCIe electrical-margin validation as part of its refurbishment/QA process, so I'm not claiming that they don't. But if the validation is primarily functional, I wonder whether a machine with a marginal PCIe path could pass refurbishment testing and only show the problem intermittently in normal use. I've raised this investigation with Intel as well as Dell, including the WHEA/AER evidence and the Lane Margining results, and I've asked Intel engineering to help interpret the processor/platform side rather than simply treating WHEA-17 as an endpoint failure. I've also supplied Intel with the Intel CrashLog/PUNIT side of the investigation. We have structurally decoded records containing MTL/NS rev9 and ARL/CDS rev3 data, including reason values 0x24 and 0x09, but we don't have verified public definitions that turn those values and the remaining payload into a diagnosis. I've therefore asked Intel to interpret the complete CrashLog/PUNIT evidence using the appropriate platform definitions, and to assess whether anything there correlates with the PCIe receiver/AER behaviour. I also approached PCI-SIG about the Lane Margining/AER side. They wouldn't provide technical support directly because that support is a membership benefit, so Intel/Dell and the upstream tooling are currently the routes I'm pursuing. One additional wrinkle: while auditing the Linux pcilmr testing I found an issue in the retained source concerning restoration of the Link Control/ASPM bits after margining. I've raised that upstream rather than ignoring it, because it potentially affects interpretation of some experimental runs. It doesn't explain the spontaneous Windows WHEA history, which predates those tests, but it does mean I'm being careful to separate test-induced/retraining observations from ordinary-use failures. That's partly why your replacement result interests me so much. If your replacement was refurbished and subsequently reproduced the same class of behaviour, I'd really like to understand what Dell actually validates on these machines before they go back into circulation—particularly whether anything equivalent to PCIe receiver-margin testing is performed. I did have the USB4 come up even though I'm not using a dock. The field engineer who was asked to exchange the port refused to as it had one system board change already. I am wondering if there is a wider manufacturer defect like the xbox360 red ring one. Have you ever run Lane Margining at the Receiver on your 285HX machine, particularly the Thunderbolt/USB4 path that reports 8086:5786? If the topology exposes margining-capable host, endpoint or retimer receivers, comparing those results could be very useful.
delboi Posted yesterday at 08:37 AM Posted yesterday at 08:37 AM (edited) 22 hours ago, whyshchuck said: Answering your hardware-swap question - sibling model here (Pro Max 18 Plus, same 285HX CPU): - System board (motherboard): replaced. No change. - Discrete NVIDIA GPU module: physically replaced. No change. - Whole unit: I was given a different unit (fresh Windows, different RAM/disk/ panel). Same behaviour returned - and two units showed an identical WHEA-17 signature side by side. - BE200 Wi-Fi: not physically removed in my case, so I can't speak to that. One difference worth noting: my sustained AER doesn't come from Wi-Fi/NVMe/GPU at boot like yours - it's from the USB4/Thunderbolt host-router downstream port (PCI\VEN_8086&DEV_5786), i.e. my dock link. So for me the BE200 doesn't look central, though the common 285HX / ARL068 angle still fits. Net: board swap, GPU-module swap and a full-unit swap all failed to fix it, which points at the platform rather than a single removable part. Quick follow-up question: as I’m located in the UK, I'm curious about the global footprint of this issue. If you're able to share, even high-level geographic or continental data would be very insightful. I recognize Dell sources localized components regionally (such as keyboard layouts), so I'm trying to determine why replacement units across different regions are exhibiting this identical failure mode. Could you share the make/model of your dock I might look into trying to get one to see if it changes my errors. I am assuming you get the same error when you run the laptop outside the dock? In an ideal RMA/refurbishment pipeline, you would expect diagnostic workflows leveraging tools like: Time Domain Reflectometry (TDR): To pinpoint physical impedance discontinuities or structural trace faults along the PCB. PCIe Gen 5 Signal Integrity Validation: A high-bandwidth real-time oscilloscope paired with a PCIe interposer (e.g., Wild River Technology) to capture eye diagrams directly off the M.2 or GPU interconnects. Protocol Analysis: A PCIe protocol analyzer (Teledyne LeCroy, Ellisys) for TLP/DLLP transaction-level validation. Given the recurring nature of this fault domain across replacement parts, my hypothesis is that these deep signal integrity and structural tests aren't being executed in the refurb loop, leading to a recirculation of defective inventory that could otherwise be reworked at the board level. Alternatively—if these are factory-new components and systems—it points toward a fundamental design fragility or layout vulnerability across the platform. For anyone interested in running a PCIe lane margining baseline test to verify link stability, here are the details and methodology: https://www.dell.com/community/en/conversations/dell-pro-max-laptops/pro-max-16-plus-mb16250-pcie-lane-margining-baseline-results/6aaa303258cc775fd9775935 This overall behavior is reminiscent of the Xbox 360 'Red Ring' saga, where the convergence of early RoHS lead-free solder transitions, underfill thermal expansion stresses during factory burn-in, and aggressive thermal cycling caused systematic BGA interconnect degradation. Edited yesterday at 09:06 AM by delboi clarifying question
delboi Posted 3 hours ago Posted 3 hours ago A status update on the WHEA-17/PCIe investigation. I would be interested to hear how Fedora 45 Beta or newer (at the time of writing comes with 7.2 kernel), or another Linux distribution, behaves for anyone here with this issue. Intel asked me to test with Linux kernel 7.2. On my Ubuntu setup, I saw no AER messages with 7.2.6, even after PCIe lane-margin testing, whereas I obtained AER reports under 7.0. The margining tool still reported failing results under the test settings used, so the quieter log does not establish that the link’s electrical margin improved. I am keeping those observations separate rather than calling 7.2 a fix. There is an important limitation: the initial tests were not a like-for-like comparison. My earlier 7.2 boot used pcie_aspm=off, and its native AER reporting enables were off, unlike the 7.0 session where AER was demonstrably reporting errors. We therefore need to verify both the power-management configuration and the reporting path before comparing results. An empty error log is inconclusive when error reporting is not operating equivalently. I also looked at Linux commit c855c992, a generic PCI power-management change present in 7.2. It concerns L1.1/L1.2 link substates—not L2—and their configuration around device sleep/wake transitions. The issue was reported with NVIDIA H100 hardware, but the change is in common PCI code. It is not documented as an ARL068 workaround, and I have not established that it explains my results. Intel lists ARL068 for the relevant Core Ultra HX processor family used by the MB16250. However, an erratum applying to a processor family is different from proving that a particular machine’s errors are caused by it. Corrected AER/WHEA events and failing margin results do not, by themselves, identify ARL068. Back in Windows, I captured a burst of 13 corrected WHEA-17 event records during a short session of manual device scans, without running lane margining. That gives us a specific activity to investigate for repeatability, rather than relying only on intermittent freezes. It does not yet establish that scanning caused the errors, identify the faulty component, or prove that Windows is responsible. I have also not established that those corrected events and the freezes share the same cause. Is anyone already running Fedora, Ubuntu or another Linux distribution on their Pro Max 16/18 Plus? I would particularly welcome results from 7.2, but other kernel versions are useful too—especially from people who also see WHEA-17 in Windows. My Ubuntu image booted 7.0.0-30-generic, and I added 7.2.6-070206-generic separately. I am interested in Fedora as a way to test a normally maintained newer kernel, not as a confirmed workaround. Most people here are Windows users, so this is not a request to reinstall or switch operating systems. Anyone already comfortable with Linux could help build a comparison. Reports from machines that do not show the problem would be useful as well. For anyone contributing, the useful details are: Machine/configuration: model, BIOS version, CPU/GPU/SSD models, SSD firmware if known, and whether you were on AC or battery or using a dock. Software/settings: distribution, exact kernel version (uname -r), Windows build for comparison, relevant driver versions or whether the GPU/Wi-Fi drivers are loaded, and any custom PCIe or power-management boot options. Observation: what you were doing, when and for how long, any AER/WHEA records or actual freezes/timeouts, and whether native AER reporting was confirmed active or has not yet been checked. The aim is to establish whether there is a repeatable difference on the same hardware, with comparable settings and working error reporting, and whether other owners see it too. That would give us a stronger test case to take back to Dell and Intel, and to submit through Microsoft Feedback Hub for investigation of the Windows behaviour. I am not claiming Linux has a workaround Microsoft is missing, that ARL068 is the root cause, or that the hardware has been cleared. I would like to turn the observations into evidence that the engineering teams can compare and investigate.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now