|
|
This topic comprises 4 pages: 1 2 3 4
|
|
Author
|
Topic: Dolby DSS 200 Array Degraded and Interrupted Playback
|
|
|
|
|
Marcel Birgelen
Film God
Posts: 3357
From: Maastricht, Limburg, Netherlands
Registered: Feb 2012
|
posted 12-13-2013 05:53 PM
quote: Scott Norwood Sometimes good disks will be reported as "failed" by the RAID controller, yet the array can safely be rebuilt without replacing them. If this happens more than once, though, it is another sign of impending failure.
That can also happen if there is buggy firmware present on your disk. Another problem are disks that aren't designed for RAID purposes. Those disks often take too long to respond to a request, especially if an automated relocation is happening, that causes the RAID controller to drop the disk from the array. While they still may operate quite stable in a software-RAID environment, they usually fail rather rapidly in a hardware RAID situation.
And then there are those absolute terror drives like the WD Green series, those that put themselves in standby mode and get kicked out of the RAID once that happens...
Personally, I'm not really fond of RAID5 for those ever increasing arrays. Rebuilds are taking longer and longer and the risk of another disk failure is especially large during a rebuild. Something like RAID6 or a software based solution like ZFS with multi-disk redundancy would probably be better suited for future storage implementations.
| IP: Logged
|
|
|
|
|
|
|
|
Steve Guttag
We forgot the crackers Gromit!!!

Posts: 12814
From: Annapolis, MD
Registered: Dec 1999
|
posted 12-14-2013 03:38 PM
There are numerous ways of remotely checking the drives and some have been listed here (e.g. a NOC with a decent SNMP system).
We have access to just about every theatre we install for the purpose of support. Lets say it is a Dolby server...how long does it take to go to the service tab and check the reallocated sector count? In GDC one can pull up the whole SMART report without ever pulling the logs. The answer by the way...seconds...you can spend about 15-seconds a screen checking that sort of stuff so you can remotely be in and out of a theatre...even a large plex in just a few minutes having checked 48 drives in a 12-plex. Again, will you catch everything? No...but it is a RAID 5 so even if a drive does go down...SNMP will already tell you that...if not the customer asking about the error/LED flashing. All without harming a show in progress.
If you see a drive with higher than normal error counts...take a quick note so the next time you go through you can compare to see how fast it is rising.
I know that as these systems are now getting older we are starting to see the effects of age/wear. One one of my regular remote checks...I found a 3-drive system were one disc had a LOT of reallocated sectors and the other two drives had a bit of a history of other ERRORs (according to the SMART report). My solution...all three drives got changed...an not one show was compromised.
As to drive types...all of the servers we work with use the Enterprise grade drives from the factory and we replace them with that sort of drive with near identical read/write speeds. Personally, I hate Seagate as seemingly all of my failures come from them (not just in the theatres...but in personal use too). I've had great luck with Hitachi and Western Digital. Again, for the WDs...it is their Enterprise grade drives, not their green energy efficient drives that are used. My Mac at home runs on one of their "Black Caviar" drives...never a lick of trouble in years and boots up real fast too. We all have our personal preferences based on past experience...my past is not your past so our opinions may vary.
I think, as time goes, we'll all develop our service techniques the same way we did for film systems...some will do better than others. I think, like with other systems, it is a mater of doing things as smartly and efficiently as possible to make the most use of one's time while minimizing problems. To me, wasting time on sending back a perfectly good drive and having the customer pay for that returned drive (the return freight is almost ALWAYS on the customer) for a drive that would have a 99.99% change of having a normal life-span just seems like a waste for all unless you could demonstrate that such a drive causes any show degradation or premature system failure. If the server manufacturer says that if a drive has 10 or more bad sectors within the warranty period they'll change it...okay...they have set that threshold. I haven't seen any manufacturer claiming just 1 bad sector is the threshold.
| IP: Logged
|
|
|
|
|
|
|
|
|
|
|
|
Leo Enticknap
Film God

Posts: 7474
From: Loma Linda, CA
Registered: Jul 2000
|
posted 02-11-2016 10:04 AM
Thought I'd write this up in case anyone else experiences this.
Arrived in the booth about an hour and a half ago, and noticed that the DSS200 was doing the same thing as it did when I had a hard drive fail a few months ago. The thing would freeze totally for about 30 seconds (as in, Show Manager totally unresponsive), then all eight hard drive green LEDs would blink for about half a second in succession, the RAID card would say "Bleep Bleep!," followed by the server unfreezing and behaving normally for about a minute, rinse and repeat.
So I immediately had a look at the theatre devices tab to see which of the drives was buggered this time (we keep a spare in the booth, so at this point I wasn't too worried), but it said that the RAID was OK, and reported no reallocated sectors on any of the drives! I tried rebooting the server (properly - entered the reboot command into the shell, then yanked both power cords after the PSU fans surged, waited for 30 seconds and then powered up again), but it still did the same thing after the reboot.
At this point I was starting to panic: we've got a big show tonight with a Hollywood A-lister in attendance, the studio's tech coming to inspect us at noon and the prospect of the DCP server throwing a hissy fit not exactly making my day. After a deep breath and a coffee, I decided to try removing each drive, cleaning the contacts and reseating them. My plan was to do the same thing with the RAID card if this didn't work. My thinking was that given that Show Manager didn't report any bad drives, the fault, if there actually was one, was probably in the RAID card.
Anyway, cleaning the contacts and reseating the drives appears to have worked. The RAID card gave a single bleep on the reboot - a hopeful sign - and it's been about an hour since then without any trouble. Furthermore it's been ingesting throughout that hour. No freezes or double bleeps so far. My guess is that repeated mechanical (ramp up and down) and/or heat cycling caused one of them to work loose, hence the RAID card's temper tantrum.
| IP: Logged
|
|
|
|
|
|
|
|
All times are Central (GMT -6:00)
|
This topic comprises 4 pages: 1 2 3 4
|
Powered by Infopop Corporation
UBB.classicTM
6.3.1.2
The Film-Tech Forums are designed for various members related to the cinema industry to express their opinions, viewpoints and testimonials on various products, services and events based upon speculation, personal knowledge and factual information through use, therefore all views represented here allow no liability upon the publishers of this web site and the owners of said views assume no liability for any ill will resulting from these postings. The posts made here are for educational as well as entertainment purposes and as such anyone viewing this portion of the website must accept these views as statements of the author of that opinion
and agrees to release the authors from any and all liability.
|