This is topic DSS-200 Degraded Array Question in forum Digital Cinema Forum at Film-Tech Forum ARCHIVE.
To visit this topic, use this URL:
https://ft-forum.com/ft/cgi-bin/ubb/ultimatebb.cgi?ubb=get_topic;f=16;t=003748
Posted by Jim Cassedy (Member # 4115) on 07-22-2019, 03:42 PM:
Today, at one of the screening rooms I work at, I noticed a "degraded array"
warning. When I investigated further, the system is showing Drive 1 with 16
reallocated sectors and Drive2 with 15 reallocated sectors. (
)
I was just at this location two days ago and I KNOW this warning was not present.
So far, it hasn't in any way affected operations.
I know they need to look at getting some new drives in there, but in the short term,
my question is: CAN I DO A RAID RE-BUILD WITH 2 DRIVES DOWN LIKE THIS?
I need to use this server again for a press screening on Thursday,
and I don't have the time to get/install new drives B4 then
Since, so far, it's not affecting operations, I'm tempted to leave it
"as is" until I can deal with it next week. (This room isn't used every day)
Posted by Mike Renlund (Member # 4675) on 07-22-2019, 04:59 PM:
Hi Jim,
Reallocated sectors are a leading indicator of drive failures, but having reallocated sectors doesn't necessarily mean that the drives are bad. I had a site with a major structural failure that shook the building very hard, and we gained a few dozen reallocated sectors added...but that was from the mechanical shock. After that they were still fine.
The DSS200 will attempt to rebuild on it's own, so at least one of those drives may be bad since it is still showing the error.
It is advisable to try and command a rebuild just in case and won't affect playback. I'd also look at the front panel and see if one or both drives show red.
If it is only one, power down, pull the drive, and power up again. If the unit still works, get a new drive for the system.
If the unit doesn't boot, you know that two drives are acting badly (you should be able to put the one drive back and limp along while you get two new drives, then wipe the system to create a new RAID). Obviously take note of settings in the config script, any serial commands and settings, and copy all the needed content off.
You can also put system logs through the Dolby Log Analyzer to see additional information.
If you need help on any of these steps, contact our team at CinemaSupport@dolby.com. We can look deeply into the logs and easily tell if you have one or two drives that are bad.
Mike Renlund
Dolby Laboratories
Posted by Leo Enticknap (Member # 534) on 07-22-2019, 08:35 PM:
I look after a few DSS200s and 220s showing between 1 and 100 reallocated sectors, one of them on all the drives, literally for years. Of course I advise replacement of any drive showing reallocated sectors as soon as I see it (whenever I touch a server, I download a log package and put it through the online analyzer, usually as the first thing I do), but sometimes, the end user doesn't regard it as a priority to do that. The server still works, right?
If the DCP you need for Thursday is ingested and you know that it plays OK, then if it were me, I'd leave the server powered up, and deal with the problem after your screening is out of the way.
Again, if it were me, I'd put a log from your server through the Dolby log analyzer. If it shows that those drives have racked up more than 43,920 hours (which is equivalent to five years of spinning: there are 8,784 hours in a year, and as a very rough rule of thumb, enterprise grade SATA drives are said to be good for five years before the risk of failure starts to increase significantly), replace the whole set and do a clean install of the software. Take the opportunity to open up and clean out the server case (e.g. with a paintbrush or Datavac), too.
2TB enterprise grade drives are now so cheap, that, IMHO, it is silly to take the risk of a RAID going out, stopping a show, and maybe having to give out 100 refunds.
Posted by Mark Gulbrandsen (Member # 72) on 07-23-2019, 08:58 AM:
quote:
2TB enterprise grade drives are now so cheap, that, IMHO, it is silly to take the risk of a RAID going out, stopping a show, and maybe having to give out 100 refunds.
I couldn't agree more with that statement!! Why risk losing a show for $300 worth of drives. The refund will cause you way more than $300 worth of grief. Average drive lifespan is three to four years anyway. Unfortunately Dolbly didn't allow for a hot swap standby drive which is really a shame since their RAID card does support it.
Mark
Posted by Jim Cassedy (Member # 4115) on 07-23-2019, 09:21 AM:
Thanks for the replies. I actually wasn't supposed to work that screening
yesterday and got called in at the very last minute & managed to squeeze
it in before another commitment afterwards. I don't know why I didn't
think of pulling & analyzing the logs, except that I was in a rush to get
to my next location. Duh!
quote: Mike Renlund
The DSS200 will attempt to rebuild on it's own,
After the screening, I tried re-booting the serverthingy, and it
automatically started doing a 'raid re-buid'. This morning, I logged in
remotely and the re-build had completed, and the "degraded array"
error message was gone.
So for now, things seem to be under control, although obviously I
need to talk to 'the boss' about getting some new drives in there.
I've installed new drives before, so I'm familiar with the procedure,
and I think there's also an instructional document on how to do it
on the Dolby Customer website, which I have access to in case I
need to refresh my memory. The Approved Replacement Drive List
is also there IIRC.
quote: Leo Enticknap
If the DCP you need for Thursday is ingested and you know that it
plays OK, then if it were me, I'd leave the server powered up, and
deal with the problem after your screening is out of the way.
Actually, the Press Screening in question was supposed to happen last
Friday, but at the last minute was re-scheduled to this Thursday, so the
content IS already ingested, and I just got a new KDM, which I was able
to (remotely) ingest this morning with no problem.
So for now, I, too believe that "If it works, don't F*** with it" (at least
until after Thursday) is the best path forward.
The room isn't being used again until then, & I don't have time to go
down there before then anyway.
- - but when I do go in on Thursday, just for fun, I'm going to pull the
logs and run them through the Dolby Log Analyzer & see what it sez.
quote: Leo Enticknap
it is silly to take the risk of a RAID going out, stopping a show, and
maybe having to give out 100 refunds
Agreed, although in our case it would only be ONE refund, albeit a big one!
Posted by Marcel Birgelen (Member # 6801) on 07-23-2019, 10:51 AM:
Keep in mind that a degraded array will affect the performance considerably. But a RAID array that's rebuilding will impact the performance even more. Also keep in mind that depending on disk size and load on the machine, a RAID rebuild can take anywhere between hours and days...
Also remember that an array that has a failed member, that has been re-inserted into array, has a big chance of failing again after a while. Often, when a disk/member fails during a show, it will have an impact on the show, like a short to medium hickup during the presentation.
Also, a faulty disk can sometimes bring the performance of your array to a grind, so much, that it even impacts normal playback. If you have a disk that's failing, but always answers back just before the RAID controller or RAID software would eject it, it can cause some severe issues with performance. Some Film-Tech members have experienced such behavior before. RAID controllers and RAID software usually isn't smart enough to detect this kind of behavior and as long as the disk keeps on giving back the requested result, it will hang on to the lagging disk.
I consider disks to be a consumable. You usually should swap them out after 3 to 5 years, depending a lot on the usage pattern on them. But for a normal server, playing one to three shows every day, I'd say replacing them every 3 years is good practice.
Luckily, compared to almost anything else on this equipment, disks are pretty cheap.
Posted by Leo Enticknap (Member # 534) on 07-23-2019, 10:56 AM:
On servers with full-sized 3.5" SATA drives that are left running 24/7, our experience is that after around 30,000 hours, the chances of bad (reallocated, as Dolby Show Manager labels them) sectors appearing increases from almost zero to low, but the chance of an outright drive failure remains near zero. After around 40 to 45,000 hours, a drive is almost guaranteed to have some bad sectors, and the chance of a drive failure becomes significant.
On IMS type servers with 2.5" drives, those figures become 20,000 and 30,000 hours. They just don't last as long, possibly because they operate at a higher temperature.
Posted by Ken Lackner (Member # 1002) on 07-23-2019, 01:19 PM:
Probably also because they are notebook drives and not enterprise grade.
Posted by Mark Gulbrandsen (Member # 72) on 07-23-2019, 04:35 PM:
I have some 2.5" enterprise grade drives running in several places. But it's too soon to tell. The cost is about 20% higher than notebook drives, and most enterprise servers run 2.5" drives now as they can fit more drives in a given server. The thing I don't get here is why are we even using mechanical drives these days? OK for the OS where it may be frequently written to, but stupid for the raid where 80% of their lives they are reading back data, and reading data is not detrimental to SSD's lifespan.
Mark
Posted by Leo Enticknap (Member # 534) on 07-23-2019, 05:15 PM:
The Ultra Stereo IMS used SSDs, and its imminent reincarnation as the Q-Sys one also does so; but that's the only example of SSDs in an IMS that I can think of. Apart from them not being approved for use under warranty coverage, I can't think of any reason why you couldn't put them in a Dolby, GDC or Barco IMS, though.
I have been told by various people that SSDs are even less tolerant of excessive heat than spinning rust drives, which I'm guessing could be why they haven't appeared much in IMS servers, as well as the cost.
It's much easier to keep old school server drives cool: in particular, the fans in a DSS200 shift enough CFM to air condition Hell. In an IMS, you're basically reliant on the card cage cooling provided in the projector.
Posted by Carsten Kurz (Member # 5396) on 07-23-2019, 08:01 PM:
Barco now also qualified SSDs for the ICMP. I think manufacturers didn't follow that path earlier because SSDs had been prohibitively expensive in usable capacities before. The two USL IMS using SSDs also supported DTS-X offline rendering and playback in addition to HFR bandwidths, so probably needed considerable higher throughput from it's RAID.
Posted by Steve Guttag (Member # 268) on 07-24-2019, 07:40 AM:
Dolby has qualified some 2TB SSD drives as well. I suspect, over time, more and more will move to SSD as the cost difference is continuing to diminish to near nothing. Reliability of SSD has also gone up.
As for 3.5" HDD, you get 5-7 years out of them, typically (unless you are into Seagate) so what's the big deal? 2.5" HDDs are also seemingly doing well. What are you going to get out of the SSD? Faster boot times? Again, I think it is all heading that way anyway.
Posted by Marcel Birgelen (Member # 6801) on 07-24-2019, 09:18 AM:
The biggest advantage of SSDs will be faster turnaround times when you can push content around over a network. Although in most cases, the network will be the new bottleneck, as most servers and IMSes are limited to gigabit speeds.
Maybe we see the hard-disk based DCP distribution swap to SSDs. Not only are SSDs far cheaper to ship (you can essentially ship them by postal mail with a pouch-slip), with a proper USB 3 or higher connection, you could enjoy a major speedup in content ingestion.
Posted by Scott Norwood (Member # 30) on 07-24-2019, 09:25 AM:
It is hard to find conclusive data on this, but mechanical hard disks appear to still be more reliable than SSDs. The write-cycle limit of SSD has effectively been solved, but they still fail (and tend to fail completely and without warning).
Given that digital cinema has a fixed maximum bit rate and that even a single 7200RPM mechanical drive can reliably read data at that rate, the advantage of SSDs in a cinema server is minimal (mostly not being sensitive to vibration and generating less heat). I could see the advantage in a library server, though.
Posted by Marcel Birgelen (Member # 6801) on 07-24-2019, 10:51 AM:
It depends on the use-case of the SSD. An SSD can theoretically read an infinite amount of data, it cannot store an infinite amount of data, because of the limited write cycle. For some applications, SSDs are more reliable than rotating rust, due to this aspect.
The write cycle limit of an SSD is pretty predictable and can be closely monitored. The problem is that, when you put them into an RAID array, all with the same age, they're going to fail awfully close together. So, you need to replace them before you park yourself against the wall.
The failure of an SSD is indeed mostly binary, it's usually completely dead. Data recovery from an SSD, if possible at all, is even more expensive than from a platter.
The biggest advantage of an SSD is the far shorter access time to random data and therefore far larger "IOPS". This performance gain is primarily experienced in stuff like databases.
For standard DCI servers, not much gain to be gotten. One advantage though is, that with SSDs you have less to worry about bandwidth problems on your storage medium when we're going to see stuff like 60fps 4K or even 120fps 4K content...
Posted by Carsten Kurz (Member # 5396) on 07-24-2019, 12:15 PM:
QSC CMS-5000 uses SSDs, and, consequently, offers a 2.5/5/10G ingest port. Up to 600MBit/s J2K decoding. Add conventional audio + ATMOS/DTS-X playout and SSDs become quite necessary for a three/four drive RAID.
Posted by Jim Cassedy (Member # 4115) on 07-24-2019, 02:02 PM:
quote: Marcel Birgelen
One advantage though is, that with SSDs you have less
to worry about bandwidth problems
Oh, geez- - I long for the days when The Studios would send me both DCP
& a 35mm print for these press screenings. I never had bandwidth problems
then, either. (and sometimes they'd even let me choose which one to play
)
Posted by Scott Norwood (Member # 30) on 07-24-2019, 05:29 PM:
Can you just play the movie off of a CRU drive if it comes that way?
(Do advance screenings get regular DCPs on CRU drives or special ones on USB drives? I've not done this sort of thing since everything came on film.)
Posted by Carsten Kurz (Member # 5396) on 07-24-2019, 06:04 PM:
Some servers allow to play DCPs right off CRU or even USB drives. It's not recommended, since you omit the content checking and validation of a full ingest, but it works. I remember that one cinema in Ireland I regularly visit plays their weekly arthouse series always directly from CRU, they don't want to spend the extra ingest time for a single screening.
- Carsten
Posted by Steve Guttag (Member # 268) on 07-24-2019, 06:16 PM:
Servers that I know can do it are any Dolby DSS server (DSS220 must use an ESATA cable, not USB). GDC servers with CRU bay (SX2100, SX2001, SX2000AR), Barco ICMP and the Dolby IMS3000.
Doremi has their "P'Ingest" but that does load the content on the server so it doesn't count (you also have to give it a 15+ minute head start).
The GDC servers I'm not sure about are the SX3000/SX4000 with an Enterprise Storage Plus. They would need two ESATA ports on each end (one for playing content off of the RAID and one for the CRU.
The IMS3000 and ICMP play off of USB3.
Posted by Mark Gulbrandsen (Member # 72) on 07-25-2019, 06:38 AM:
SX-3000 has three eSATA ports, but only one can be selected at any given time and the USB is really nbot quite a USB-3.But the current CRU bays are eSATA, USB-3 and FW800. SO the bay is at least capable of direct playback. I see CRU's actually going away pretty soon for the most part anyway. Satellite delivery is quickly becomming the norm, even with my customers.
Mark
Posted by Gordon McLeod (Member # 33) on 07-25-2019, 06:40 AM:
The SX3000 will play off a cru drive either with the Enterprise Plus or with a mobile one plugged into the Esata port
Posted by Jim Cassedy (Member # 4115) on 07-25-2019, 09:15 AM:
quote: Scott Norwood
Can you just play the movie off of a CRU drive if it comes that way?
(Do advance screenings get regular DCPs on CRU drives or special
ones on USB drives?
What I'm mostly involved with are press & industry screenings which
are not open to the general public, and in fact from which the public
is sometimes specifically excluded. (So if a journalist or sound-tech,
or someone brings a guest who is not a card carrying member of the
press or SAG/AFTRA/DGA or IATSE, they don't get in)
But to answer your question, Scott, 95% of the material arrives on
standard CRU drives, although the drives & KDMS are sometimes
not booked or shipped through standard distribution channels but
come from some publicity arm or the studio instead.
Sometimes the drives I get were produced before the ones that
are mass produced for theatrical distribution, so they arrive with
hand-written labels, or labels different than the ones that will go
to the theaters. The drives may also contain a pre-show slide or
even a short intro from the film-maker that the studio or publicity
company wants us to put on screen before the show.
I get most of these drives in plenty of time to ingest before a
screening, but sometimes because they are still 'tweaking' things
in post production or for some other reason the film-maker or
publicity agent will hand-carry it to the event, and I will have to
do a 'live-play' if I don't have time to ingest. I make sure the
client is aware of the risks of doing that, but I've never had any
glitches while doing a live play, usually from a DSS-200.
(Perhaps it's partly because most of these drives I get are
usually brand new & not "recycled" ones?)
Some exceptions to the 'standard ingest & playback' have been:
> One of the larger places work has direct hi-speed fiber connections
to some of the studios down in LA, and I've done one or two screenings
where they were literally still working post on post production the night
before the press event, and so we got the flick by high-speed transfer.
> Also, and again because we were doing a screening where the movie
was still undergoing post production, we did the screenings direct from
the DCDM 'work print'. The DCDM is on a standard CRU drive, and the
server it plays on looks exactly like a DSS-200, except the face plate
is a totally different color. (And of course there are internal technical
differences, since you can't play a DCDM DCP on a standard server)
> For one screening that was booked at the last minute, the film-maker
realized he had brought the wrong version of the DCP with him and did
not want to screen that one. But he did have the correct version on
an HDCAM tape with Dolby E sound with him and I wound up running
the show off of that.
Posted by Mark Gulbrandsen (Member # 72) on 07-25-2019, 11:22 AM:
quote: Gordon McLeod
The SX3000 will play off a cru drive either with the Enterprise Plus or with a mobile one plugged into the Esata port
Yes, as long as you have a later eSATA eqipped CRU bay. Early on, not all bays had eSATA, just USB and regular SATA. Now for some reason they have added Firewire 800. Talk about adding an obsolete connection system...
Mark
Posted by Greg Routenburg (Member # 1742) on 07-25-2019, 03:37 PM:
Does anyone know if there's an official reason why Doremi/Dolby has never incorporated a "Live Play" feature on their ShowVault-IMB line of servers? The P-Ingest thing is pretty useless if you're trying to save a show if the content arrives through the door at the last minute.
Posted by Steve Kraus (Member # 476) on 07-25-2019, 05:29 PM:
If you were going to replace all four drives is it better replace all at once, reload software, and then content? Or replace one at a time, letting it rebuild each time. Does it work out the same in the end?
Posted by Leo Enticknap (Member # 534) on 07-25-2019, 05:47 PM:
Replace all four drives, do a clean install of the software with the current version (4.9.1.22, assuming that, if you're using the cat862, the projector accepts TLS link encryption), and if you're using a cat862, a "reinstall all components" reinstallation of the media block firmware, too. Take the opportunity to give the case a good cleanout, replace the CMOS battery, and check the fans for any worn bearings, while you're at it. You will then effectively have a new server.
Of course it's important to make a careful note of any configuration settings (ctrl+alt+F1, then login as config) and Show Manager settings, including automation cues, before you proceed, so you can reconstruct them on the fresh installation. You'll also want to "outgest" any content that you do not have available for reingestion on external drives.
RAID drives are a bit like car tires - I would always suggest replacing the complete set at once if you can, so that wear is even (unlike car tires, though, you don't need to rotate them!). Otherwise, you're worrying about some of the drives having done fewer hours than others, which can be a pain to keep track of.
Posted by Steve Guttag (Member # 268) on 07-25-2019, 10:25 PM:
To a degree, Greg, I suspect that Doremi's issue with "live-play" was probably baked into the system from the get-go and without a work around on how to stream the data with the available hardware.
P'ingest works for many as you only need a 15-minute head start so if the drive walks in the door, stick it in and you are only waiting 15 minutes or so instead of the full ingest duration. I'm not defending it and I don't consider that to be an actual live-play (because it isn't).
By far the most flexible on that are the Dolby DSS servers...play off the drive, transfer content at full speed to the RAID...whatever you want to do.
I haven't played with the IMS3000 that much but the ICMP seems pretty robust on its Play Now feature. USB3.0 seems plenty fast enough. It even works on USB2 though I suspect if I had a longer piece, there would have been underflows.
Posted by Mark Gulbrandsen (Member # 72) on 07-26-2019, 03:22 PM:
quote: Steve Guttag
By far the most flexible on that are the Dolby DSS servers...play off the drive, transfer content at full speed to the RAID...whatever you want to do.
Not quite.... No eSATA.
Posted by Steve Guttag (Member # 268) on 07-26-2019, 03:54 PM:
No need...they have a CRU bay. The DSS220 doesn't have a CRU bay but DOES have eSATA on front and rear for external.
Posted by Mark Gulbrandsen (Member # 72) on 07-27-2019, 10:33 AM:
quote: Scott Norwood
It is hard to find conclusive data on this, but mechanical hard disks appear to still be more reliable than SSDs. The write-cycle limit of SSD has effectively been solved, but they still fail (and tend to fail completely and without warning).
I have never had an SSD completely fail over the last 15 years of using them. I have only lost data here and there on one drive too, and those bad sectors get placed in the "do not use bin" by the drive firmware. Fact is, if you have an electronic failure with either SSD or mechanical drive, the drive is all done. So I don't get why we are not using SSD's.
Mark
Posted by Marcel Birgelen (Member # 6801) on 07-27-2019, 11:50 AM:
For many applications I prefer SSDs, simply because of the speed. For a workstation, going back to hard drives feels like going back to the stone age.
But if you look at a NAS for example, the situation is not that clear cut. It really depends on the type of data on there and the access patterns.
The primary reason for not using SSDs in cinema servers would be the cost difference. The gap between "consumer grade" SSDs and hard drives is still steadily closing in regarding to pricing and "consumer grade" SSDs should be sufficient for most DCI applications.
Please notice that even those SSDs that are sold as "enterprise" are often just the "consumer" version and not the datacenter version, which is often more than 4 times as expensive. The difference between both is usually the write speed and the number of supported write cycles.
Regarding failure modes: I've had several SSDs go completely bad, they where in a state where nothing worked. I've also had "half-failures" in a RAID, where one defective SSD would make the RAID array crawl. It was pretty hard to find the SSD causing it, because there were like 12 SSDs in the array and none reported any problems.
With hard drives it's almost never a black and white situation, you usually get some warning signs before a hard-drive fails. Even if the electronics fail on a hard drive, data recovery companies are usually capable of still extracting the information from there. Only if a drive arm buries itself into the platter and turns the surface to dust, it's usually over-and-out with the data on there. If the controller of an SSD blows up, there is often no way of getting your data back. SSDs store their data nowhere near sequentially, this is done to balance usage of an SSD more evenly. When the table where what data is stored gets corrupted, it's bye bye for your data. I've seen the latter even happening after firmware updates, where the format of this index changed and the new firmware corrupted the table... I believe it was OCZ branded drive that this happened to.
Posted by Steve Guttag (Member # 268) on 07-28-2019, 05:28 AM:
I just installed a new IMS2000 and the drives were all 2TB Micron SSDs, from the factory.
According to the latest approved drive sheet, the only approved 2TB drives for the IMS2000 and IMS3000 are Micron SSDs. (two models of them).
Powered by Infopop Corporation
UBB.classicTM
6.3.1.2