Jump to content
NotebookTalk

All Activity

This stream auto-updates

  1. Past hour
  2. Hi guys, My RTX 3000 suddenly shut the laptop off while gaming. Now the charger LED goes out immediately when I plug it into the laptop with the RTX installed. I tested with battery removed and i had the same issue I removed the RTX and the charger light stays on I plugged in my old 765m in and charger works normally So it looks like the fault is on the RTX 3000 itself. I don’t see any obvious burnt components. Could it be a shorted mosfet or capacitor or vrm component, and is this usually repairable?
  3. Today
  4. A status update on the WHEA-17/PCIe investigation. I would be interested to hear how Fedora 45 Beta or newer (at the time of writing comes with 7.2 kernel), or another Linux distribution, behaves for anyone here with this issue. Intel asked me to test with Linux kernel 7.2. On my Ubuntu setup, I saw no AER messages with 7.2.6, even after PCIe lane-margin testing, whereas I obtained AER reports under 7.0. The margining tool still reported failing results under the test settings used, so the quieter log does not establish that the link’s electrical margin improved. I am keeping those observations separate rather than calling 7.2 a fix. There is an important limitation: the initial tests were not a like-for-like comparison. My earlier 7.2 boot used pcie_aspm=off, and its native AER reporting enables were off, unlike the 7.0 session where AER was demonstrably reporting errors. We therefore need to verify both the power-management configuration and the reporting path before comparing results. An empty error log is inconclusive when error reporting is not operating equivalently. I also looked at Linux commit c855c992, a generic PCI power-management change present in 7.2. It concerns L1.1/L1.2 link substates—not L2—and their configuration around device sleep/wake transitions. The issue was reported with NVIDIA H100 hardware, but the change is in common PCI code. It is not documented as an ARL068 workaround, and I have not established that it explains my results. Intel lists ARL068 for the relevant Core Ultra HX processor family used by the MB16250. However, an erratum applying to a processor family is different from proving that a particular machine’s errors are caused by it. Corrected AER/WHEA events and failing margin results do not, by themselves, identify ARL068. Back in Windows, I captured a burst of 13 corrected WHEA-17 event records during a short session of manual device scans, without running lane margining. That gives us a specific activity to investigate for repeatability, rather than relying only on intermittent freezes. It does not yet establish that scanning caused the errors, identify the faulty component, or prove that Windows is responsible. I have also not established that those corrected events and the freezes share the same cause. Is anyone already running Fedora, Ubuntu or another Linux distribution on their Pro Max 16/18 Plus? I would particularly welcome results from 7.2, but other kernel versions are useful too—especially from people who also see WHEA-17 in Windows. My Ubuntu image booted 7.0.0-30-generic, and I added 7.2.6-070206-generic separately. I am interested in Fedora as a way to test a normally maintained newer kernel, not as a confirmed workaround. Most people here are Windows users, so this is not a request to reinstall or switch operating systems. Anyone already comfortable with Linux could help build a comparison. Reports from machines that do not show the problem would be useful as well. For anyone contributing, the useful details are: Machine/configuration: model, BIOS version, CPU/GPU/SSD models, SSD firmware if known, and whether you were on AC or battery or using a dock. Software/settings: distribution, exact kernel version (uname -r), Windows build for comparison, relevant driver versions or whether the GPU/Wi-Fi drivers are loaded, and any custom PCIe or power-management boot options. Observation: what you were doing, when and for how long, any AER/WHEA records or actual freezes/timeouts, and whether native AER reporting was confirmed active or has not yet been checked. The aim is to establish whether there is a repeatable difference on the same hardware, with comparable settings and working error reporting, and whether other owners see it too. That would give us a stronger test case to take back to Dell and Intel, and to submit through Microsoft Feedback Hub for investigation of the Windows behaviour. I am not claiming Linux has a workaround Microsoft is missing, that ARL068 is the root cause, or that the hardware has been cleared. I would like to turn the observations into evidence that the engineering teams can compare and investigate.
  5. Try with Secure Boot disabled as well. EDIT: In addition, disabling the Above 4GB MMIO BIOS assignment while enabling above 4g decoding should fix the PCI resources error. Resizable BAR option is not visible on my P775TM1.
  6. Do you have above 4g decoding and resizable bar turned on ? In the bios . Might actually fix the issue related to ram .
  7. No you cannot use standard MXM 3.0 b or 3.1b cards in the clevo laptops that have separate power connections . The slots are wired differently
  8. These will not work without a wire mod . Cisco uses proprietary MXM cards ( the wires I forget which ones are reversed but they are related to power or ground ) that's why your cards don't work . You need two clevo / MSI MXM cards with the power connectors . Be aware 3080s draw up to 125 watt . So make sure your cooking is sufficient m I have a single 3080ti in mine . It's a 200 watt monster of a GPU i have a GTX 1060 in there just for emergency in case the 3080 exercise croaks . The laptop will auto switch to pcie MXM 1
  9. I replaced a K5100. The old pads were just fine. They don't degrade, unless you've had water in there, or had it apart before and torn them up. I wouldn't worry about it. Do clean off and replace the grease though, every time that surface separates.
  10. Iam replacing old K3100, i bet these pads are quite not good, i want to have a good heat transfer on the P5000, i was looking at 0.2mm high PTM7950 from Honeywell, they should be alright, right?
  11. I replaced nVidia with nVidia, and reused the same original pads. Seems to work just fine. If you're replacing AMD with nVidia, then I wouldn't know.
  12. I was able to change the edp output but i have no way of testing it with the quadro p4200 cause i brick my m6700s mxm slot when i was adjusting the thermal pads. i copy the code section from the quadro p4000 engineerng vbios to a quadro p4200 oc imac vbios. anyone be interested of trying out this vbios for edp output? P4200_R2test3
  13. Yesterday
  14. Congratulations, but how long did the test take? 🙂 A couple of minutes?.. If so, that’s not quite right - these tests are often configured to check only 10–20% of the total video memory (or a specific fixed small amount). That might be enough in case of a death of single memory chip, but for more complex issues, I strongly recommend to let the GPU warm up well during the long full test. To enable a full test, you need to open the MATS configuration file named 'runmats' for editing - and locate the line at the very end that calls the MATS program; it look something like this: "$LOCATION/$PKGNAME/mats" -e 20 (check only 20 Mb of VRAM) or "$LOCATION/$PKGNAME/mats" -c 10 (check only 10 percent of VRAM) ...and delete any symbols after 'mats"'. 🙂 Or you may comment non-modified line with symbols '#' and a following ' ', both placed in the beginning, and a add new one with a call of MATS without arguments. And you may wish to add as the last line of this file the command: poweroff Then your PC will be auto powered off shortly after end of check (it's very useful in case of long test caused huge size of VRAM). The results will be written into text .log file placed in the same directory. Enjoy. 😉
  15. That’s the problem—I can't do it; I press the key, but absolutely nothing happens. I read somewhere that the "USB Legacy Support" setting in the BIOS might be disabled. I remember being able to do that when I had UEFI with CSM (legacy) support enabled, but now it's full UEFI—and possibly has USB legacy support disabled too—yet the system doesn't detect the external keyboard during POST. The external keyboard works fine in Windows, of course, but pressing F2 won't let me enter the BIOS. Regarding the TDP limit: I had a similar issue with a 120W MSI GTX 1070. To avoid constantly toggling settings—since I had to go through that whole "sg -> igfx -> sg" routine almost every other time I booted up—I just set the limit to 100W in MSI Afterburner (I don't even recall if I flashed the BIOS to unlock that slider in the Mobile Pascal TDP Editor, but it doesn't matter). It seems like the same thing is happening here—maybe it can't draw more power because of that? Or is it a Max-Q BIOS, which is why the power draw is capped at 90W?
  16. Update: i ran the mods MATS test, and seems like memories still good, it was a flawless PASS; still need to put the card under charge to see if it works normally; and ow i did flash a new VBIOS on it (MSI techpowerup) via NVFLASHK, was a success and the card is working fine, but i think i will need to do some additional tests to be sure of everything Little Update on PSU side: for the little story, i did buy a 230W PSU for the workstation, but it was not recognize by the Laptop, because it did not have the special Pin that allow the PSU to charge the laptop. So had to mod it, with the use of 100K ohm resitor to close the circuit, cut the charging cable from the original HP PSU, solder it on the new generic one (V+ & GND), and solder the Pin (central tiny cable) with a resistor to the V+ cable of the generic PSU, and just like that, i have a working 230W PSU working fine. Maybe it can help someone here, we never know.
  17. Iam set on getting the P5000 for my M6800, only picture i seen with new thermal pads is this one using AMD heatsink, which pads should i order for the nvidia heatsink? how tall should they be?
  18. You kann use an external Keyboard to enter the bios. I think the SG to IGFX Trick was for wen you have the 40w TDP issue, i hat to do it evrye time i use my M17x r4 whit an 980m becouse it woud Switch frome Boost to idle will an Benchmark
  19. Quick follow-up question: as I’m located in the UK, I'm curious about the global footprint of this issue. If you're able to share, even high-level geographic or continental data would be very insightful. I recognize Dell sources localized components regionally (such as keyboard layouts), so I'm trying to determine why replacement units across different regions are exhibiting this identical failure mode. Could you share the make/model of your dock I might look into trying to get one to see if it changes my errors. I am assuming you get the same error when you run the laptop outside the dock? In an ideal RMA/refurbishment pipeline, you would expect diagnostic workflows leveraging tools like: Time Domain Reflectometry (TDR): To pinpoint physical impedance discontinuities or structural trace faults along the PCB. PCIe Gen 5 Signal Integrity Validation: A high-bandwidth real-time oscilloscope paired with a PCIe interposer (e.g., Wild River Technology) to capture eye diagrams directly off the M.2 or GPU interconnects. Protocol Analysis: A PCIe protocol analyzer (Teledyne LeCroy, Ellisys) for TLP/DLLP transaction-level validation. Given the recurring nature of this fault domain across replacement parts, my hypothesis is that these deep signal integrity and structural tests aren't being executed in the refurb loop, leading to a recirculation of defective inventory that could otherwise be reworked at the board level. Alternatively—if these are factory-new components and systems—it points toward a fundamental design fragility or layout vulnerability across the platform. For anyone interested in running a PCIe lane margining baseline test to verify link stability, here are the details and methodology: https://www.dell.com/community/en/conversations/dell-pro-max-laptops/pro-max-16-plus-mb16250-pcie-lane-margining-baseline-results/6aaa303258cc775fd9775935 This overall behavior is reminiscent of the Xbox 360 'Red Ring' saga, where the convergence of early RoHS lead-free solder transitions, underfill thermal expansion stresses during factory burn-in, and aggressive thermal cycling caused systematic BGA interconnect degradation.
  20. Last week
  21. Thanks for the response I am starting to think this is a fault domain which a PCIe signal integrity issue somewhere between: The remaining fault domains are approximately: CPU PCIe Root Complex / Root Port hardware — including the host-side PCIe PHY/receiver. Motherboard PCIe channel — traces, connectors and other physical elements between the Root Port and endpoint. Power delivery affecting the PCIe subsystem. PCIe reference clock / clocking behaviour. Firmware/software-controlled PCIe behaviour — UEFI/BIOS, ACPI/platform power management, link-state management, link training/retraining and related platform configuration. Your full-unit replacement result is particularly interesting to me because Dell is currently discussing replacing my MB16250 with a Certified Refurbished unit. Was your original machine new or refurbished, and was the complete replacement Dell supplied new or refurbished? If refurbished, do you know anything about its previous repair history? One concern I have is what Dell's refurbishment validation actually tests at PCIe level. I have been using PCIe Lane Margining at the Receiver (LMR) under Linux as part of my investigation. My understanding is that this interface exists specifically to assess receiver timing/voltage margin rather than merely checking that a PCIe device enumerates, negotiates its expected link speed/width and passes ordinary functional tests. I don't know whether Dell performs LMR or equivalent PCIe electrical-margin validation as part of its refurbishment/QA process, so I'm not claiming that they don't. But if the validation is primarily functional, I wonder whether a machine with a marginal PCIe path could pass refurbishment testing and only show the problem intermittently in normal use. I've raised this investigation with Intel as well as Dell, including the WHEA/AER evidence and the Lane Margining results, and I've asked Intel engineering to help interpret the processor/platform side rather than simply treating WHEA-17 as an endpoint failure. I've also supplied Intel with the Intel CrashLog/PUNIT side of the investigation. We have structurally decoded records containing MTL/NS rev9 and ARL/CDS rev3 data, including reason values 0x24 and 0x09, but we don't have verified public definitions that turn those values and the remaining payload into a diagnosis. I've therefore asked Intel to interpret the complete CrashLog/PUNIT evidence using the appropriate platform definitions, and to assess whether anything there correlates with the PCIe receiver/AER behaviour. I also approached PCI-SIG about the Lane Margining/AER side. They wouldn't provide technical support directly because that support is a membership benefit, so Intel/Dell and the upstream tooling are currently the routes I'm pursuing. One additional wrinkle: while auditing the Linux pcilmr testing I found an issue in the retained source concerning restoration of the Link Control/ASPM bits after margining. I've raised that upstream rather than ignoring it, because it potentially affects interpretation of some experimental runs. It doesn't explain the spontaneous Windows WHEA history, which predates those tests, but it does mean I'm being careful to separate test-induced/retraining observations from ordinary-use failures. That's partly why your replacement result interests me so much. If your replacement was refurbished and subsequently reproduced the same class of behaviour, I'd really like to understand what Dell actually validates on these machines before they go back into circulation—particularly whether anything equivalent to PCIe receiver-margin testing is performed. I did have the USB4 come up even though I'm not using a dock. The field engineer who was asked to exchange the port refused to as it had one system board change already. I am wondering if there is a wider manufacturer defect like the xbox360 red ring one. Have you ever run Lane Margining at the Receiver on your 285HX machine, particularly the Thunderbolt/USB4 path that reports 8086:5786? If the topology exposes margining-capable host, endpoint or retimer receivers, comparing those results could be very useful.
  22. I haven't heard anything about newer drivers other than mvolt+ not working with some driver versions.
  23. Hey guys! I have recently gotten back An Alienware 15 r3 that I Gifted my cousin years ago. He's now upgraded and Gave it me back, The laptop devolved some issues with the age of the device So I need to order some parts for it. The battery Puffed up and damaged the keyboard, Palmrest and chassis cover. It would cost a bit to buy the parts to get it pristine again. It also has some dead Blue LEDS on the Alienware lid logo and "alienware" on the front bezel. I have always wanted to convert one of these into an 17 R4 so I am tempted to at least research and attempt it. Has anyone tried this before and had any problems? Or is it just a straight drop in with the motherboard? As I would just buy the parts to convert it. The specs of the 15 r3: I7 7700HQ 32GB DDR3 2133Mhz GTX 1070M 8GB 120Hz TN Display I am choosing to do this to save another alienware from landfill. I use a x15 R1 as my main laptop. So anyone with any Guidance Or experience Please let me know! Thanks for reading!
  24. Guys, I remember someone here posting that the new 6xx drivers are not good. Can someone catch me up on why? I'm on 596.xx on both desktop & laptops. I finally got SLI working on 596.xx drivers using the same old patch for 446.xx drivers. SO now I am aiming for 616.xx drivers.
  25. Hi everyone. This is an old topic, but I still have both of my M18x R2s, and I’ve run out of working keyboards...) None of them have fully functional keys anymore—except for one where the 'N' and '-' keys still work. But I need the F2 key. Does anyone have any ideas or experience on how to enter the BIOS with a completely non-functional keyboard? I have full UEFI enabled (no CSM). I recall that with UEFI with CSM support—back when the big alien logo appeared at startup—I could press a key on an external keyboard to enter the BIOS, but now I'm on full UEFI. I just need to get in once. By the way, I recently installed a Quadro RTX 5000 for a good price—I was thinking of posting about it in a separate thread, but I can mention here that it draws a maximum of 85–90 watts. I seem to remember there’s a way to switch from SG mode to iGFX and back to get full power. Are there any other options? Wasn't there a trick involving sleep mode? I can't seem to find or recall the details.
  26. M5 Ultra vs a 9950X3D + 5090 kitted out PC: CPU: Bonkers, next level CPU performance. Absolutely trashes Intel and AMD at 150w with their 36 core model both single and multi. Single core: ~44% faster Multi core: ~80% faster All this while consuming 25% less power in a small, compact box that makes zero noise 99% of the time and even under load is a fraction of the noise a typical PC generates. You're going to need a Threadripper 9970x (32 core) to compete, but not beat, at this point on multi. To flat out beat Apple, you will need a 9980x. On both, they absolutely get destroyed for single core performance. On Single, Intel and AMD are going to have to bring some serious magic in their next gen CPUs or Apple will just continue to dominate and I don't think they're going to even get close to catch Apple in single threaded performance. Looking forward to the 52 core Intel Nova "Open all salvos!" Lake beast vs the M5 Ultra. Heck, even the 28 core might get close in multi. Blender Render = M5 Ultra is ~72% faster GPU: Time to come back to reality Apple vs Nvidia..... Optimized for Nvidia, 5090 pulls ahead showing Nvidia's strength in terms of raw power, vertical integration and software stacks. LLMs: If memory isn't an issue, 5090 is a beast but obviously 32GB can't stack up against 96GB or 256GB. Toss in an RTX 6000 (or two or three) and it's game over but then again you can score a fully kitted M5 Ultra with 96GB of memory and 1TB storage and max CPU/GPU config for $6799.99 for less than half the price of an RTX 6000. You can score a 256GB model for $10799 which is 2/3 the cost of a single RTX 6000 at this point. In the end, Apple's massively upgraded Neural Processors are no match for Nvidia's raw power but their unified memory architecture is where they have always shined. Expect the 512GB model (not available yet) to cost ~$17k as the 256GB model is currently $10,799.,99 and upgrading from 96GB to 256GB is a $4k up charge. So you end up with a 512GB M5 Ultra for what is quickly becoming the cost of a single RTX 6000. Video editing, M5 Ultra: DaVinci = tied Abobe Pro = M5 Ultra finished in less than half the time Gaming: Time for a brutal reality check. Even Metal optimized, 5090 just trashes the M5 Ultra and is over twice as fast. M5 Ultra looks to be right between 5070ti-5080 in gaming performance but looking at some of the results that are actually taking advantage of the 80 core GPU, there is some optimizations to be had but....yeah, the 5090 just drops the hammer on Apple....just like Nvidia dropped the hammer on AMD with the 5090. My final verdict? M5 CPU architecture is next level. Absolutely next level. GPU has really come along way even vs M4 and its strength lies in its unified memory bandwidth especially in these AI driven days. For my personal use, target LLMs for particular use cases the 5090's 32GB is working out just fine and this is where it absolutely shines. If you need more than 32GB, performance scarily drops off as expected and you need to start looking at the 5080 based 4500 card with 48GB or even 5500 (barf) up to a 6000 to match the 96GB config from Apple. 256GB config for $10,799 really makes a compelling config for many small AI shops, startups or individuals because it costs less than a single 6000 by almost $5k and you get over 2.5x the memory. I suspect the 512GB will also be a very compelling buy for those on a budget but needing 512GB of memory.
  27. Answering your hardware-swap question - sibling model here (Pro Max 18 Plus, same 285HX CPU): - System board (motherboard): replaced. No change. - Discrete NVIDIA GPU module: physically replaced. No change. - Whole unit: I was given a different unit (fresh Windows, different RAM/disk/ panel). Same behaviour returned - and two units showed an identical WHEA-17 signature side by side. - BE200 Wi-Fi: not physically removed in my case, so I can't speak to that. One difference worth noting: my sustained AER doesn't come from Wi-Fi/NVMe/GPU at boot like yours - it's from the USB4/Thunderbolt host-router downstream port (PCI\VEN_8086&DEV_5786), i.e. my dock link. So for me the BE200 doesn't look central, though the common 285HX / ARL068 angle still fits. Net: board swap, GPU-module swap and a full-unit swap all failed to fix it, which points at the platform rather than a single removable part.
  28. this is why i now have a junker dual core 8570W (only difference is no RAM slot 1&3 under keyboard from what i can see) in a half chassis for my heatsink work, all i need are CPU and GPU bolt points to align. i can prep all i need without killing anything (fingers crossed) plus i can mod the chassis for cooling and swap items across as proven. no spare 980m though (eek)
  29. Hi, I’ve been trying to find a replacement for my PC for some time, but I’m having trouble finding the right laptop because every high-end model seems to have some kind of drawback. There are a few things that are particularly important to me: I want a display with low response times, as I mainly play fast-paced FPS games such as Apex Legends, Deadlock, CS, and Valorant. I want good thermals and, ideally, no CPU or GPU thermal throttling. It would also be preferable if the WASD area didn’t get excessively hot during gaming. I specifically want the Ryzen 9 9955HX3D, as it seems to be significantly faster in gaming than the Intel Core Ultra 9 275HX. The laptop should have at least an RTX 5080. I’ve already looked at several models: Lenovo Legion 7 Pro – On paper, this laptop looks great, but the CPU temperatures are honestly quite concerning. 100°C while gaming? And these weren’t even synthetic benchmarks such as Cinebench — it reached those temperatures while playing Cyberpunk 2077... ASUS ROG Strix SCAR 18 – This laptop initially looked very promising, but its display seems to suffer from some kind of strange flickering caused by a conflict between the backlight strobing and synchronization technology. You can see the effect in Linus’ video here: https://youtu.be/mpDanDwd7x0?t=168. Apparently, this can potentially cause headaches during longer gaming sessions. I also looked at laptops with external liquid cooling from Dream Machines and XMG. However, it seems that these models are available with displays that have relatively slow response times, making them less suitable for fast-paced games due to noticeable ghosting. Are there any other laptops currently on the market that I might have overlooked that could meet these requirements?
  1. Load more activity
×
×
  • Create New...

Important Information

We have placed cookies on your device to help make this website better. You can adjust your cookie settings, otherwise we'll assume you're okay to continue. Terms of Use