Film-Tech Cinema Systems
Film-Tech Forum ARCHIVE


  
my profile | my password | search | faq & rules | forum home
  next oldest topic   next newest topic
» Film-Tech Forum ARCHIVE   » Operations   » Digital Cinema Forum   » Dolby DSS 200 Array Degraded and Interrupted Playback (Page 3)

 
This topic comprises 4 pages: 1  2  3  4 
 
Author Topic: Dolby DSS 200 Array Degraded and Interrupted Playback
Marco Giustini
Film God

Posts: 2713
From: Reading, UK
Registered: Nov 2007


 - posted 12-13-2013 05:07 PM      Profile for Marco Giustini   Email Marco Giustini   Send New Private Message       Edit/Delete Post 
Steve,

I did not mean to criticise your or anybody's service procedures, if it looked like that I apologise of course.

Your point is solid but I still believe in mine. I don't know how I could regularly access every system and track the number of bad sectors on every single drive. It means to track how the bad sectors number behaves on thousands of HDDs.

That being said, if I know the projectionist is a good one I could simply ask them to check the status of a 'suspected' drive over time and send a replacement if/when it's needed.
But the catch on the previous sentence is "is a good one" [Smile]

Chris,
It's very unlikely a bad sector can be repaired. And as I said time is money. I would definitely try to repair a disk in my home computer but when it comes to professional service where the customer is paying for my time and expecting a fully up and running system when I leave, that does not apply. IMHO of course.

 |  IP: Logged

Scott Norwood
Film God

Posts: 8146
From: Boston, MA. USA (1774.21 miles northeast of Dallas)
Registered: Jun 99


 - posted 12-13-2013 05:26 PM      Profile for Scott Norwood   Author's Homepage   Email Scott Norwood   Send New Private Message       Edit/Delete Post 
Quite a bit of my day job involves storage issues, and I am with Brad and Steve on this: all disks have bad sectors. This is not a problem per se, but a rapid increase in the reallocated sector count (SMART parameter) does correlate well with impending disk failure (see the Google study on disk failure).

Also, I would pay attention to the hardware ECC recovered parameter (if there is a way to view this).

That said, SMART is not actually all that good at detecting disks that are about to fail. Really, we shouldn't have to worry all that much about this stuff as long as the RAID controller is doing its job. Sometimes good disks will be reported as "failed" by the RAID controller, yet the array can safely be rebuilt without replacing them. If this happens more than once, though, it is another sign of impending failure.

 |  IP: Logged

Marcel Birgelen
Film God

Posts: 3357
From: Maastricht, Limburg, Netherlands
Registered: Feb 2012


 - posted 12-13-2013 05:53 PM      Profile for Marcel Birgelen   Email Marcel Birgelen   Send New Private Message       Edit/Delete Post 
quote: Scott Norwood
Sometimes good disks will be reported as "failed" by the RAID controller, yet the array can safely be rebuilt without replacing them. If this happens more than once, though, it is another sign of impending failure.
That can also happen if there is buggy firmware present on your disk. Another problem are disks that aren't designed for RAID purposes. Those disks often take too long to respond to a request, especially if an automated relocation is happening, that causes the RAID controller to drop the disk from the array. While they still may operate quite stable in a software-RAID environment, they usually fail rather rapidly in a hardware RAID situation.

And then there are those absolute terror drives like the WD Green series, those that put themselves in standby mode and get kicked out of the RAID once that happens...

Personally, I'm not really fond of RAID5 for those ever increasing arrays. Rebuilds are taking longer and longer and the risk of another disk failure is especially large during a rebuild. Something like RAID6 or a software based solution like ZFS with multi-disk redundancy would probably be better suited for future storage implementations.

 |  IP: Logged

Chris Slycord
Film God

Posts: 2986
From: 청주시, 경북도, South Korea
Registered: Mar 2007


 - posted 12-13-2013 06:06 PM      Profile for Chris Slycord   Email Chris Slycord   Send New Private Message       Edit/Delete Post 
On whether you can/can't keep track of the number of bad sectors on thousands of disks, am I wrong in assuming that one could run something like the linux smartd daemon to track the count day to day on each server then send an email to you if a particular disk starts changing relatively quickly?

 |  IP: Logged

Marco Giustini
Film God

Posts: 2713
From: Reading, UK
Registered: Nov 2007


 - posted 12-13-2013 06:16 PM      Profile for Marco Giustini   Email Marco Giustini   Send New Private Message       Edit/Delete Post 
Scott,

Don't get me wrong, what you say is sound.
I don't know you but my idea of average cinema is a forsaken booth managed by people who barely know how to press play. Yes, you can find all the information you need in the logs but it takes some time to download a set of logs and cannot be done while the server is playing.

Sure, you can access the server by terminal and cat/pico the files themselves but I find it a little impractical?

Marcel,
Again, you can do that on your PC or on servers you build. I would never run anything on a D-Cinema server which does not come from the manufacturer.

 |  IP: Logged

Carsten Kurz
Film God

Posts: 4340
From: Cologne, NRW, Germany
Registered: Aug 2009


 - posted 12-13-2013 08:21 PM      Profile for Carsten Kurz   Email Carsten Kurz   Send New Private Message       Edit/Delete Post 
quote: Marco Giustini
It means to track how the bad sectors number behaves on thousands of HDDs.
That would be necessary now, yes, but of course should be done by the servers - detect a sudden increase in bad sectors and report automatically. Doremis recent reporting is a clear improvement. At some time, it may also consider reporting potentially dangerous RAID conditions. You could, however, also request these from the logs or via SNMP automatically. The reason it is not done now is that a single drive failure is already taken care of by the RAID redundancy, which then SHOULD trigger an alert.
Besides that, Google demonstrated that SMART reports are not a too reliable indicator of future drive failures.

- Carsten

 |  IP: Logged

Steve Guttag
We forgot the crackers Gromit!!!

Posts: 12814
From: Annapolis, MD
Registered: Dec 1999


 - posted 12-14-2013 03:38 PM      Profile for Steve Guttag   Email Steve Guttag   Send New Private Message       Edit/Delete Post 
There are numerous ways of remotely checking the drives and some have been listed here (e.g. a NOC with a decent SNMP system).

We have access to just about every theatre we install for the purpose of support. Lets say it is a Dolby server...how long does it take to go to the service tab and check the reallocated sector count? In GDC one can pull up the whole SMART report without ever pulling the logs. The answer by the way...seconds...you can spend about 15-seconds a screen checking that sort of stuff so you can remotely be in and out of a theatre...even a large plex in just a few minutes having checked 48 drives in a 12-plex. Again, will you catch everything? No...but it is a RAID 5 so even if a drive does go down...SNMP will already tell you that...if not the customer asking about the error/LED flashing. All without harming a show in progress.

If you see a drive with higher than normal error counts...take a quick note so the next time you go through you can compare to see how fast it is rising.

I know that as these systems are now getting older we are starting to see the effects of age/wear. One one of my regular remote checks...I found a 3-drive system were one disc had a LOT of reallocated sectors and the other two drives had a bit of a history of other ERRORs (according to the SMART report). My solution...all three drives got changed...an not one show was compromised.

As to drive types...all of the servers we work with use the Enterprise grade drives from the factory and we replace them with that sort of drive with near identical read/write speeds. Personally, I hate Seagate as seemingly all of my failures come from them (not just in the theatres...but in personal use too). I've had great luck with Hitachi and Western Digital. Again, for the WDs...it is their Enterprise grade drives, not their green energy efficient drives that are used. My Mac at home runs on one of their "Black Caviar" drives...never a lick of trouble in years and boots up real fast too. We all have our personal preferences based on past experience...my past is not your past so our opinions may vary.

I think, as time goes, we'll all develop our service techniques the same way we did for film systems...some will do better than others. I think, like with other systems, it is a mater of doing things as smartly and efficiently as possible to make the most use of one's time while minimizing problems. To me, wasting time on sending back a perfectly good drive and having the customer pay for that returned drive (the return freight is almost ALWAYS on the customer) for a drive that would have a 99.99% change of having a normal life-span just seems like a waste for all unless you could demonstrate that such a drive causes any show degradation or premature system failure. If the server manufacturer says that if a drive has 10 or more bad sectors within the warranty period they'll change it...okay...they have set that threshold. I haven't seen any manufacturer claiming just 1 bad sector is the threshold.

 |  IP: Logged

Brad Miller
Administrator

Posts: 17775
From: Plano, TX (36.2 miles NW of Rockwall)
Registered: May 99


 - posted 12-14-2013 05:39 PM      Profile for Brad Miller   Author's Homepage   Email Brad Miller       Edit/Delete Post 
For giggles, we have a large complex with 4 screens off on it's own wing we lovingly call "death row", as those auditoriums only seat a handful of people and exist purely for move-down booking purposes.

They have DSS200 servers in them and for kicks I tested my theory before they opened by putting a mix and match of drives in them, all of them used. One enterprise drive, one consumer drive, one green drive, etc...all different specs and manufacturers. Some were even repaired discs back from the manufacturer. We couldn't make them fail so we left them there and kept a real close eye on them.

They have been open for over a year now without a single failure. Gotta love a real hardware raid!

 |  IP: Logged

Steve Guttag
We forgot the crackers Gromit!!!

Posts: 12814
From: Annapolis, MD
Registered: Dec 1999


 - posted 12-14-2013 07:03 PM      Profile for Steve Guttag   Email Steve Guttag   Send New Private Message       Edit/Delete Post 
That is what makes you giggle? You really need to get out more [Razz]

 |  IP: Logged

Marcel Birgelen
Film God

Posts: 3357
From: Maastricht, Limburg, Netherlands
Registered: Feb 2012


 - posted 12-14-2013 07:52 PM      Profile for Marcel Birgelen   Email Marcel Birgelen   Send New Private Message       Edit/Delete Post 
quote: Marco Giustini
Again, you can do that on your PC or on servers you build. I would never run anything on a D-Cinema server which does not come from the manufacturer.
It was more a general observation, I would not advice anybody to mess with an existing configuration, if that would be even possible. Since most current servers only carry 3 or 4 disks, it would even be quite useless.

But there is obviously a trend where the "traditional" function of the server is progressively being transferred into the IMB and the server is becoming a storage-only affair.

Just like it happened in many data centers, you will slowly see a demand for centralized storage, that's not only being used for distribution, but for direct playback as well. Such a storage solution is not something that I would ever consider running off a RAID5 system. If that's a desirable thing is something that can be debated of course.

 |  IP: Logged

Marco Giustini
Film God

Posts: 2713
From: Reading, UK
Registered: Nov 2007


 - posted 12-15-2013 08:28 AM      Profile for Marco Giustini   Email Marco Giustini   Send New Private Message       Edit/Delete Post 
In fact today I ended up with a server showing 118 bad sectors on it. Using GREP under Linux I was able to see the previous RAID status from the logs in a blink. The same disk was showing 0 bad sectors just a few days ago, then 4-10-17-20 over the following days.

 |  IP: Logged

System Notices
Forum Watchdog / Soup Nazi

Posts: 215

Registered: Apr 2004


 - posted 02-11-2016 10:04 AM      Profile for System Notices         Edit/Delete Post 

It has been 788 days since the last post.


 |  IP: Logged

Leo Enticknap
Film God

Posts: 7474
From: Loma Linda, CA
Registered: Jul 2000


 - posted 02-11-2016 10:04 AM      Profile for Leo Enticknap   Author's Homepage   Email Leo Enticknap   Send New Private Message       Edit/Delete Post 
Thought I'd write this up in case anyone else experiences this.

Arrived in the booth about an hour and a half ago, and noticed that the DSS200 was doing the same thing as it did when I had a hard drive fail a few months ago. The thing would freeze totally for about 30 seconds (as in, Show Manager totally unresponsive), then all eight hard drive green LEDs would blink for about half a second in succession, the RAID card would say "Bleep Bleep!," followed by the server unfreezing and behaving normally for about a minute, rinse and repeat.

So I immediately had a look at the theatre devices tab to see which of the drives was buggered this time (we keep a spare in the booth, so at this point I wasn't too worried), but it said that the RAID was OK, and reported no reallocated sectors on any of the drives! I tried rebooting the server (properly - entered the reboot command into the shell, then yanked both power cords after the PSU fans surged, waited for 30 seconds and then powered up again), but it still did the same thing after the reboot.

At this point I was starting to panic: we've got a big show tonight with a Hollywood A-lister in attendance, the studio's tech coming to inspect us at noon and the prospect of the DCP server throwing a hissy fit not exactly making my day. After a deep breath and a coffee, I decided to try removing each drive, cleaning the contacts and reseating them. My plan was to do the same thing with the RAID card if this didn't work. My thinking was that given that Show Manager didn't report any bad drives, the fault, if there actually was one, was probably in the RAID card.

Anyway, cleaning the contacts and reseating the drives appears to have worked. The RAID card gave a single bleep on the reboot - a hopeful sign - and it's been about an hour since then without any trouble. Furthermore it's been ingesting throughout that hour. No freezes or double bleeps so far. My guess is that repeated mechanical (ramp up and down) and/or heat cycling caused one of them to work loose, hence the RAID card's temper tantrum.

 |  IP: Logged

Tony Bandiera Jr
Film God

Posts: 3067
From: Moreland Idaho
Registered: Apr 2004


 - posted 02-11-2016 11:23 AM      Profile for Tony Bandiera Jr   Email Tony Bandiera Jr   Send New Private Message       Edit/Delete Post 
quote:
Anyway, cleaning the contacts and reseating the drives appears to have worked. The RAID card gave a single bleep on the reboot - a hopeful sign - and it's been about an hour since then without any trouble. Furthermore it's been ingesting throughout that hour. No freezes or double bleeps so far. My guess is that repeated mechanical (ramp up and down) and/or heat cycling caused one of them to work loose, hence the RAID card's temper tantrum.
Wait, so are you saying you shut down the server when not in use? The part I put in bold seems to suggest so.

There was a very long (and contentious) thread on whether to leave servers on 24/7 or shut them down....consensus was that since server hardware is designed to be on 24/7 it should stay on for maximum hard drive life. (And so far history has shown that to be correct.)

If you are not doing so already, leave the server ON 24/7 and if you haven't already, invest in a good quality UPS for the server and projector electronics. (Especially important if you are doing screenings for A-List folks.)

 |  IP: Logged

Leo Enticknap
Film God

Posts: 7474
From: Loma Linda, CA
Registered: Jul 2000


 - posted 02-11-2016 11:29 AM      Profile for Leo Enticknap   Author's Homepage   Email Leo Enticknap   Send New Private Message       Edit/Delete Post 
No, we're leaving the server on 24/7, but the hard drives won't be spinning all the time, will they? Surely the RAID card will spin them down after a given period of time in which nothing is asking to read or write to the array? That was what I had in mind by mechanical and/or heat cycling.

 |  IP: Logged



All times are Central (GMT -6:00)
This topic comprises 4 pages: 1  2  3  4 
 
   Close Topic    Move Topic    Delete Topic    next oldest topic   next newest topic
 - Printer-friendly view of this topic
Hop To:



Powered by Infopop Corporation
UBB.classicTM 6.3.1.2

The Film-Tech Forums are designed for various members related to the cinema industry to express their opinions, viewpoints and testimonials on various products, services and events based upon speculation, personal knowledge and factual information through use, therefore all views represented here allow no liability upon the publishers of this web site and the owners of said views assume no liability for any ill will resulting from these postings. The posts made here are for educational as well as entertainment purposes and as such anyone viewing this portion of the website must accept these views as statements of the author of that opinion and agrees to release the authors from any and all liability.

© 1999-2020 Film-Tech Cinema Systems, LLC. All rights reserved.