Jump to content
NotebookTalk

delboi

Member
  • Posts

    5
  • Joined

  • Last visited

Everything posted by delboi

  1. A status update on the WHEA-17/PCIe investigation. I would be interested to hear how Fedora 45 Beta or newer (at the time of writing comes with 7.2 kernel), or another Linux distribution, behaves for anyone here with this issue. Intel asked me to test with Linux kernel 7.2. On my Ubuntu setup, I saw no AER messages with 7.2.6, even after PCIe lane-margin testing, whereas I obtained AER reports under 7.0. The margining tool still reported failing results under the test settings used, so the quieter log does not establish that the link’s electrical margin improved. I am keeping those observations separate rather than calling 7.2 a fix. There is an important limitation: the initial tests were not a like-for-like comparison. My earlier 7.2 boot used pcie_aspm=off, and its native AER reporting enables were off, unlike the 7.0 session where AER was demonstrably reporting errors. We therefore need to verify both the power-management configuration and the reporting path before comparing results. An empty error log is inconclusive when error reporting is not operating equivalently. I also looked at Linux commit c855c992, a generic PCI power-management change present in 7.2. It concerns L1.1/L1.2 link substates—not L2—and their configuration around device sleep/wake transitions. The issue was reported with NVIDIA H100 hardware, but the change is in common PCI code. It is not documented as an ARL068 workaround, and I have not established that it explains my results. Intel lists ARL068 for the relevant Core Ultra HX processor family used by the MB16250. However, an erratum applying to a processor family is different from proving that a particular machine’s errors are caused by it. Corrected AER/WHEA events and failing margin results do not, by themselves, identify ARL068. Back in Windows, I captured a burst of 13 corrected WHEA-17 event records during a short session of manual device scans, without running lane margining. That gives us a specific activity to investigate for repeatability, rather than relying only on intermittent freezes. It does not yet establish that scanning caused the errors, identify the faulty component, or prove that Windows is responsible. I have also not established that those corrected events and the freezes share the same cause. Is anyone already running Fedora, Ubuntu or another Linux distribution on their Pro Max 16/18 Plus? I would particularly welcome results from 7.2, but other kernel versions are useful too—especially from people who also see WHEA-17 in Windows. My Ubuntu image booted 7.0.0-30-generic, and I added 7.2.6-070206-generic separately. I am interested in Fedora as a way to test a normally maintained newer kernel, not as a confirmed workaround. Most people here are Windows users, so this is not a request to reinstall or switch operating systems. Anyone already comfortable with Linux could help build a comparison. Reports from machines that do not show the problem would be useful as well. For anyone contributing, the useful details are: Machine/configuration: model, BIOS version, CPU/GPU/SSD models, SSD firmware if known, and whether you were on AC or battery or using a dock. Software/settings: distribution, exact kernel version (uname -r), Windows build for comparison, relevant driver versions or whether the GPU/Wi-Fi drivers are loaded, and any custom PCIe or power-management boot options. Observation: what you were doing, when and for how long, any AER/WHEA records or actual freezes/timeouts, and whether native AER reporting was confirmed active or has not yet been checked. The aim is to establish whether there is a repeatable difference on the same hardware, with comparable settings and working error reporting, and whether other owners see it too. That would give us a stronger test case to take back to Dell and Intel, and to submit through Microsoft Feedback Hub for investigation of the Windows behaviour. I am not claiming Linux has a workaround Microsoft is missing, that ARL068 is the root cause, or that the hardware has been cleared. I would like to turn the observations into evidence that the engineering teams can compare and investigate.
  2. Quick follow-up question: as I’m located in the UK, I'm curious about the global footprint of this issue. If you're able to share, even high-level geographic or continental data would be very insightful. I recognize Dell sources localized components regionally (such as keyboard layouts), so I'm trying to determine why replacement units across different regions are exhibiting this identical failure mode. Could you share the make/model of your dock I might look into trying to get one to see if it changes my errors. I am assuming you get the same error when you run the laptop outside the dock? In an ideal RMA/refurbishment pipeline, you would expect diagnostic workflows leveraging tools like: Time Domain Reflectometry (TDR): To pinpoint physical impedance discontinuities or structural trace faults along the PCB. PCIe Gen 5 Signal Integrity Validation: A high-bandwidth real-time oscilloscope paired with a PCIe interposer (e.g., Wild River Technology) to capture eye diagrams directly off the M.2 or GPU interconnects. Protocol Analysis: A PCIe protocol analyzer (Teledyne LeCroy, Ellisys) for TLP/DLLP transaction-level validation. Given the recurring nature of this fault domain across replacement parts, my hypothesis is that these deep signal integrity and structural tests aren't being executed in the refurb loop, leading to a recirculation of defective inventory that could otherwise be reworked at the board level. Alternatively—if these are factory-new components and systems—it points toward a fundamental design fragility or layout vulnerability across the platform. For anyone interested in running a PCIe lane margining baseline test to verify link stability, here are the details and methodology: https://www.dell.com/community/en/conversations/dell-pro-max-laptops/pro-max-16-plus-mb16250-pcie-lane-margining-baseline-results/6aaa303258cc775fd9775935 This overall behavior is reminiscent of the Xbox 360 'Red Ring' saga, where the convergence of early RoHS lead-free solder transitions, underfill thermal expansion stresses during factory burn-in, and aggressive thermal cycling caused systematic BGA interconnect degradation.
  3. Thanks for the response I am starting to think this is a fault domain which a PCIe signal integrity issue somewhere between: The remaining fault domains are approximately: CPU PCIe Root Complex / Root Port hardware — including the host-side PCIe PHY/receiver. Motherboard PCIe channel — traces, connectors and other physical elements between the Root Port and endpoint. Power delivery affecting the PCIe subsystem. PCIe reference clock / clocking behaviour. Firmware/software-controlled PCIe behaviour — UEFI/BIOS, ACPI/platform power management, link-state management, link training/retraining and related platform configuration. Your full-unit replacement result is particularly interesting to me because Dell is currently discussing replacing my MB16250 with a Certified Refurbished unit. Was your original machine new or refurbished, and was the complete replacement Dell supplied new or refurbished? If refurbished, do you know anything about its previous repair history? One concern I have is what Dell's refurbishment validation actually tests at PCIe level. I have been using PCIe Lane Margining at the Receiver (LMR) under Linux as part of my investigation. My understanding is that this interface exists specifically to assess receiver timing/voltage margin rather than merely checking that a PCIe device enumerates, negotiates its expected link speed/width and passes ordinary functional tests. I don't know whether Dell performs LMR or equivalent PCIe electrical-margin validation as part of its refurbishment/QA process, so I'm not claiming that they don't. But if the validation is primarily functional, I wonder whether a machine with a marginal PCIe path could pass refurbishment testing and only show the problem intermittently in normal use. I've raised this investigation with Intel as well as Dell, including the WHEA/AER evidence and the Lane Margining results, and I've asked Intel engineering to help interpret the processor/platform side rather than simply treating WHEA-17 as an endpoint failure. I've also supplied Intel with the Intel CrashLog/PUNIT side of the investigation. We have structurally decoded records containing MTL/NS rev9 and ARL/CDS rev3 data, including reason values 0x24 and 0x09, but we don't have verified public definitions that turn those values and the remaining payload into a diagnosis. I've therefore asked Intel to interpret the complete CrashLog/PUNIT evidence using the appropriate platform definitions, and to assess whether anything there correlates with the PCIe receiver/AER behaviour. I also approached PCI-SIG about the Lane Margining/AER side. They wouldn't provide technical support directly because that support is a membership benefit, so Intel/Dell and the upstream tooling are currently the routes I'm pursuing. One additional wrinkle: while auditing the Linux pcilmr testing I found an issue in the retained source concerning restoration of the Link Control/ASPM bits after margining. I've raised that upstream rather than ignoring it, because it potentially affects interpretation of some experimental runs. It doesn't explain the spontaneous Windows WHEA history, which predates those tests, but it does mean I'm being careful to separate test-induced/retraining observations from ordinary-use failures. That's partly why your replacement result interests me so much. If your replacement was refurbished and subsequently reproduced the same class of behaviour, I'd really like to understand what Dell actually validates on these machines before they go back into circulation—particularly whether anything equivalent to PCIe receiver-margin testing is performed. I did have the USB4 come up even though I'm not using a dock. The field engineer who was asked to exchange the port refused to as it had one system board change already. I am wondering if there is a wider manufacturer defect like the xbox360 red ring one. Have you ever run Lane Margining at the Receiver on your 285HX machine, particularly the Thunderbolt/USB4 path that reports 8086:5786? If the topology exposes margining-capable host, endpoint or retimer receivers, comparing those results could be very useful.
  4. I’m investigating this further with the recurrent WHEA-Logger Event 17 / PCIe AER events on a Dell Pro Max 16 Plus MB16250 with Core Ultra 9 285HX. They occur during early Windows boot across multiple unrelated PCIe endpoints: Intel BE200 Wi-Fi, NVMe SSD and NVIDIA GPU. Raw WHEA/CPER records show the same AER signature: UncorrectableErrorStatus = 0x100000 = Unsupported Request CorrectableErrorStatus = 0xa000 = Advisory Non-Fatal + Header Log Overflow Captured TLP headers decode as PCIe Configuration Read Type 0 requests, apparently probing absent functions/BDFs during early PCI enumeration. Some requests occur only microseconds apart. Linux PCI enumeration source: https://github.com/torvalds/linux/blob/master/drivers/pci/probe.c Linux PCIe AER documentation: https://www.kernel.org/doc/html/latest/PCI/pcieaer-howto.html Intel Arrow Lake Series 2 specification update / errata: https://edc.intel.com/content/www/us/en/design/products/platforms/details/arrow-lake-s/core-ultra-200s-series-processors-specification-update/errata-details/ Relevant Intel erratum: ARL068 — PCIe Gen5 Link Exit from L1 Sub-state Low Power State. It specifically mentions short back-to-back PCIe configuration-space accesses as one trigger condition. Dell also has an official article stating that boot-time WHEA-Logger ID17 is “expected behavior as per Intel” on certain earlier Intel HX systems. Note that Dell’s listed affected systems are 12th-gen/HX platforms, not the MB16250/285HX, so it does not by itself prove this newer system is expected to behave the same way: https://www.dell.com/support/kbdoc/en-uk/000216115/laptops-with-12th-gen-and-12th-gen-hx-intel-core-processors-may-display-warning-message-whea-loggerid17 WHEA-17 itself is quite generic, and I think the associated symptoms may be more significant than the event alone. In my case, I’ve seen intermittent freezing and occasional mouse slowdowns. Dell support has also told me that some users experience BSODs or system crashes. What I’m really looking for here is other people’s ideas, observations and possible avenues to investigate
  5. I’m following this thread because I’m investigating similar WHEA-Logger Event ID 17 behaviour on a Pro Max 16 Plus. One question I have for people who have had Dell work on affected systems: were the removable PCIe devices ever physically isolated or substituted during troubleshooting? In particular: was the NVIDIA DGFF module ever replaced with a known-good module, rather than only changing drivers/settings or replacing the system board? was the Intel BE200 ever physically removed/replaced, rather than only disabled in Windows or BIOS? I’m asking because I’ve seen reports involving the same BE200 PCI ID (8086:272B) on some other recent systems, including Alienware and Acer machines, and I’m interested in whether anyone has compared those cases or performed a true hardware swap test. I’m not suggesting the BE200 is necessarily responsible. For context, Dell has already replaced the system board and NVMe SSD in my system and the behaviour remains. I also ran an Ubuntu Live comparison. Linux enabled PCIe AER and the BE200, NVMe and Intel graphics were present, but I did not observe the same Windows-style AER burst during that session. The NVIDIA GPU was detected but was not fully exercised under the Linux live environment, so I don’t regard that test as proving or excluding any particular component. I’d be interested to hear specifically from anyone who has had either the BE200 or DGFF GPU physically swapped or removed and whether that changed the WHEA behaviour or the intermittent stuttering.
×
×
  • Create New...

Important Information

We have placed cookies on your device to help make this website better. You can adjust your cookie settings, otherwise we'll assume you're okay to continue. Terms of Use