2020年9月15日 星期二

Intel Xeon Platinum 8269CY 處理器使用及兼容性報告

Intel Xeon Platinum 8269CY 這款 CPU 是“阿里雲” (Alibaba Cloud) 的定制版。在超微 Supermicro X11DPi-NT 主板上 (BIOS 3.1a)測試成功。不過,該 CPU 僅支援 2666MHz 的内存(2933MHz 的也運行於這個速度),這也是一個要注意的地方。

該處理器支援 Intel Speed Select 技術(型號中帶 'Y' 的 SKU 都有的),有兩個 Profile (預設方案):標準的 26 核,運行於 2.5 GHz 的基礎速度,全核運行最高速度 3.2 GHz;以及高性能的 20 核,運行於 3.1 GHz 的基礎速度,全核運行最高速度 3.5 GHz。兩個方案的單核最高速度均爲 3.8 GHz,兩個方案的 TDP 都是 205W。方案可以在 BIOS 裏面設定。


 

按照耗電量來説,除非這個定制版還有什麽特殊的能力,否則其功耗  205W 有點過高。不過勝於這個是步進 7 的正式版 CPU,穩定性和兼容性和其他零售版的 Platinum Xeon 相若,價錢合適的話也是個不錯的入場選項。

2020年8月22日 星期六

Farewell: Hard Disk Drives

It has been so many years already - my first HDD was a CONNER CFS-420A with capacity 420 MB. It should be around 1994. That was the exciting moment I can still remember - what an upgrade to 1.44MB floppy disks!

More than 25 years later, now I have made the decision to phase out all HDD sin my system. I have even bought a few 12TB Seagate Exos for my next project, but now they will be sold before being put into real use.

My experiences with HDDs could be quite a bit different from most of the people out there: I was very cautious about reliability and durability, rather than focusing mainly on capacity. If you have ever noticed about the specifications of the most decent HDDs, you might start noticing one thing: the Unrecoverable Error Rate has remained 10^14 (consumer parts) and 10^15 (server parts for many years without improvement.

Yes, right, who would be really caring about these numbers? They are just indicators, and new technology should be always more reliable, right? Unfortunately, my personal believe is, if there is a number published in that way, it MUST has a meaning. Consider other common storage technologies, such as SSD (10^17 for enterprise parts) and Tapes (10^17 - 10^19), what I can say is HDDs have a frightening LOW reliability nowadays.

Imagine a 1 TB HDD and a 10 TB HDD, they both have URE of 10^15. During a RAID rebuild, 10TB one will have 10 times higher chance to encounter an URE. While 10^15 URE means you *might* encounter a Non-Recoverable Read Error every 125 TB of data, this is not a lot for today's large capacity drives. If I am using 10x 12TB Exos HDDs to build a RAID 6 array, it will be so likely to get a Read Error during re-build - which is exactly the reason putting me off from proceeding with my original plan. I just don't want to put my data at risk.

That's also the reason why in the past my 16 HDDs RAID 60 array served me well, because each member was only 1TB in size. With 10^15 URE, the array was still relatively secured and it has proved itself. However, with 12TB, 10^15 URE is DANGEROUS and it is not making sense for traditional RAID techonology.

Also, think about the rebuild time - with my 16 HDDs RAID 60, even with a 1TB drive size, the rebuild time was somewhere around 4 hours without anyone using it. How about a 12 TB HDD? Well, if you are luck, you can have the array rebuilt in 2 days (given that nobody will be using it). Otherwise, it could be "weeks" - and the longer the time, the more likely another unrecoverable read error can happen. (And another one!)

You may say RAID is now out of date. But unfortunately it is the opposite - HDDs are now out of date. I now have an array of 2TB SSDs in RAID 6, and I am feeling way more comfortable with their 10^17 URE. The maximum size of SSD I can accept to be put into the array would be 15.36TB - if higher then I will need 10^18 URE.

HDD has a good characteristic: offline durability. This is especially true when it comes to offline data storage - think about data retention of SSD when it has been completely powered off - after the Intel 320 array incident happened to me in the past (data gone after powering off for > 1 month), I think you might expect data could be held in server class SSDs for 1 year max offline without problem, but don't expect more than that. HDD can do far better than SSD in this case since the storage of data is not depending on electricity stored in NAND cells that could leak, but on platter with magnetic recording that can last.

HOWEVER, how often are you going to put your HDDs offline? Possibly never - they are in RAID, and they are serving as nearline storage. This has defeated the purpose of this great characteristic. For nearline storage, I can easily use SSDs because they are ONLINE, with continuous power supply so they have the same data retention reliability as HDDs. They can run cooler and they have far better access performance than HDDs. So that's why my 12TB Exos are out of the picture for my project - I will be getting some 15.36 TB SSDs instead.

For archiving, an "old school" technology has somehow came to my mind - TAPES! They are offline storage, portable, power efficient and has a really high URE. They use magnetic recording as well so data retention is surprisingly good. If I need archiving, a tape library is better suited in this case.

With Tapes and SSDs, I have to say "farewell" to HDDs - unless URE has been increased to 10^16 or even 10^17, they won't be considered by me just because I want to have better sleep at night without worrying about RAID rebuild...

FAREWELL, MY LOVELY HDDS!

2020年8月1日 星期六

Windows Network Direct: Your better bet is with Windows Server 2019

I have been always struggling to get RDMA working inside a Windows virtual machine. I had tried Mellanox ConnectX-3, Mellanox ConnectX-5 and Chelsio T62100-LP-CR network adapters, with Windows Server 2012 or 2016, and even with Direct Device Assignment in Windows Server 2016, I could not get RDMA working flawlessly in any virtual machine.

Recently, I retried RDMA in a Windows VM (2019) on a Windows Server 2019 host, with a Chelsio T62100-LP-CR - and finally have RDMA (iWarp) working correctly (even without a switch - you can connect port 1 to port 2 to form a 100GbE link). It enabled SMB Direct between the VM and the host (and between VMs as well), and performance was acceptable (needs tuning).

If you are after any RDMA application inside a Windows VM, or simple just want to use SMB Direct in a VM, Windows Server 2019 or later is your better bet in this case.

Do note that you need the following:
  • SR-IOV support from BIOS - this sometimes means enabling the ASPM option in BIOS.
  • A network card that supports RDMA - I like iWarp because it is simpler (virtually no configuration needed). If you like RoCE then you may need DCB configured properly, or even need a 40GbE/100GbE switch.
  • Windows Server 2019 or higher - both host and VM. You may use Windows 10 (latest) - I didn't try that out but theoretically it should work.
  • Workstation / Server grade hardware - I have seen many times people complaining about not being able to enable SR-IOV due to missing implementation like IOMMU or ACS etc. with consumer grade hardware. Your CPU supports all these features doesn't mean your motherboard / BIOS has support of all features.

2016年8月28日 星期日

Windows Server 2016 DDA 直接設備授予虛擬機的一些要點

算是微軟的一些“突破”吧,反正 VMWare 一早就有這個技術。不論如何,至少直接把 PCI Express 設備,例如顯卡,聲卡,網卡或者 USB 3.0 端口授予虛擬機單獨訪問已經變得可能。然而并非所有設備都都能夠使用此功能。掃描了一下自己伺服器的硬件大概折這樣的:

1、老式 PCI 設備——別想了,放棄吧。
2、PCI Express 1.0 設備或者鏈路到 PCI Express 1.0 橋接芯片上的設備——也別想了,洗洗睡吧。
3、連路到 Intel PCH (就是現在有點想南橋的那個東東) 的設備——看運氣,就算你是 PCI Express 2.0,沒有 ACS (Access Control Service) 那也是別想了……什麽意思呢?就是板載的網卡,USB 3.0 口之類的,衹有是從 PCH 延伸出來的,就別想了……
4、有 ACS,但是是複合設備,那也別想了,至少我的 NVIDIA Quadro K4200 是不行的!

所以最終發現能夠 DDA 的就衹有直接連在 CPU 的根部 PCI Express 3.0 端口的設備,或者鏈路到支持 ACS 的 PCI Express 2.0 擴展盒的設備:
- Dell PERC H730P (別開玩笑了……把 100 個磁盤弄到 VM 裏面幹啥?)
- Mellanox ConnectX-3 (好吧,算是靠譜一點……)
- HP ioDrive DUO SLC 320GB (算不算 NVMe 直接授予!?-_-)

看來我那些 PCI Express 1.0 的 Expansion Backplanes / Enclosures 都沒啥用處了~

NVIDIA Grid 授權模式:要買蘋果得把高壓鍋也買了!?

我并不反對 NVIDIA 對 GRID 的授權模式的更改。儘管每用戶的授權會增加花費,但是從價格上還是可以接受。但是!但是!最主要的是,就算你入手了一張 NVIDIA GRID M10,如果沒有從 NVIDIA 指定合作夥伴購買整機,你是沒有辦法獲得授權的!

什麽意思?就是你從 ebay 買了十張 NVIDIA GDID M10 卡那也衹是一堆廢件!除非你從 Dell 或者 Supermicro 什麽的整機購買了伺服器。這樣你才能夠從 NVIDIA 購買授權……

簡單點:想買蘋果?先把這個高壓鍋買了 。怎麽?你不用蘋果來燒湯?那我可不管,買高壓鍋可是先決條件呢……

2015年12月24日 星期四

微軟已經修復 Windows Server 2016 TP 中的 Deduplication 數據損壞問題

近日收到 Microsoft 的消息說 Windows Server 2016 TP 中的 Deduplication 數據損壞問題已經被修復,并且附上了一個内部測試補丁。安裝該補丁之後,測試運行了幾次 DedupJob 均沒有問題。產生問題的是 dedup.sys 驅動程序。

看來微軟對於數據損壞此類嚴重問題還是比較重視的。如果閣下希望在 Windows Server 2016 TP 中使用 Deduplication,則需要等待補丁在 Windows Update 中的正式發佈。比較保險的應該會在下一個纍積更新之中,或者更保險一點可以等待至 TP5。注意必須先安裝補丁,然后再開啓 Deduplication 功能。從微軟工程師的描述中看來此問題可能衹會發生在從 Windows 2012 R2 升級到 2016 的系統中。

關於詳細問題描述可以參考此帖子

2015年12月11日 星期五

[已修復] 注意 Windows Server 2016 TP 中的 Deduplication 可能會導致數據損壞

UPDATE: 此問題已修復

不知道這算是幸運還是不幸,反正就被我遇上了。基本情況如下:

先決條件:
- 系統是 Windows 2012 R2
- 磁盤爲 GPT 的 NTFS,開啓 VDI 模式的 Deduplication,數據重複刪除率達到 75%
- 狀態: 1TB 中刪除重複后大約使用 250GB。
- 磁盤儲存大量 Windows 2008 R2 與 Windows 8.1 的 VM,格式爲 VHD 或者  VHDX
- 磁盤是本地磁盤,注意這個配置 Microsoft 不建議。沒有 SAN 或者 iSCSI 的使用。

步驟:
- 升級 Host 到 Windows 2016 TP4,并且安裝 Deduplication 功能
- 將所有 VM 導入到 Hyper-V,并且運行
- 添加更多的 VM
- 確保 Background Deduplication 運行至少一次

結局:
- 大部分 VM 突然進入 BSOD 狀態
- 檢查該 VM 的 VHD / VHDX 文件,發現無法用 CHKDSK 修復,數據完全丟失。卷返回 Invalid Function 錯誤。
- 該 VHD / VHDX 文件無法重複使用!你必須刪除該文件,然後重新創建,才能夠在 VM 中重新安裝系統
- Host 中開啓 Deduplication 的卷卻沒有問題,CHKDSK 通過。
- 關閉 Background Deduplication 后,就不會進一步損壞其他數據

所以此次數據損壞可能是 Dedup 服務造成的。已經將此問題報告 Microsoft,他們也在進一步調查,不過在他們回復之前,閣下最好還是先關閉 Dedup 服務以避免產生同樣的問題。

Incompatibilities and Compatibilities

NOTE: This article will be updated in the future when more compatibilities / incompatibilities are discovered.  Incompatibilities   12-Feb-...