Random losses of communication between DSS220 and cat745 IMB

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Marco Giustini
    Film God
    • Jan 2020
    • 1168
    • Reading, UK

    #31
    thinking about it, it makes sense. The HW RAID in the 200 was activated at BIOS level hence the array was online before the OS was booting.

    That said, even my old PC had a software RAID which was still controlled by the motherboard BIOS - hence it was available immediately at boot time. I am not an expert on this matter so I am not sure how it worked. It just feels so weird that losing disk 1 was killing the system.
    That said, even nowadays the OS sits on small "SSDs" on newer servers so it's not unusual.

    But yes, this won't help Leo

    Comment

    • Bruce Cloutier
      Pro Film Handler
      • Jan 2020
      • 429
      • Pittsburgh, PA USA

      #32
      Originally posted by Caleb Williams
      I strongly encourage the at minimum weekly reboot for the DSS220. It's a little hacky because of the union-fs but you can create a crontab that will do the reboot automatically for you directly on the server if there's not another option.
      I am not a fan of periodic rebooting. That is a work-around rather than a solution. Someone at the company thinks "Yeah just reboot. That is good. The customer is happy. They go away." Engineers at said company are isolated from it. They think "There are no major problems. We do good." But you know there is someone there thinking "Wait! They have to reboot all of the time? Our stuff sucks."

      I liken it to flying an aircraft. The risky operations are takeoff and landing. Those are analogous to boot and shutdown. That system reboot doesn't eliminate the original problem. It just postpones it. So if you are into procrastination, go for it. But worse, You don't really know if the shutdown really is controlled and safe. It likely leaves things half done. Especially if there are bugs at the root of it. Then on reboot will it come back up fully? Is it back just as it was before?

      You guys are phobic about updates. Well, a reboot carries some of the same risk. I would be just a little concerned about that.

      When I hear about my stuff being rebooted. I cringe. We research why? And, we do something about it. That doesn't mean that anyone stopped rebooting. Or, even updated to the fixed firmware. Nope.

      Meanwhile, apparently, even our contacts at Barco have gone stale. We're working on it.


      Comment

      • Marco Giustini
        Film God
        • Jan 2020
        • 1168
        • Reading, UK

        #33
        indeed - I feel it's important to make very clear that a reboot is not an inevitable task because software inevitably goes wrong over time. Proper SW and proper HW don't go wrong. A proper device would be able to work indefinitely. Bits are bits. 0's and 1's won't change.

        A reboot means "the software has messed up the system so much we need to start from scratch". But it will eventually go back into that situation again.

        Comment

        • Leo Enticknap
          Film God
          • Jan 2020
          • 3572
          • Loma Linda, CA

          #34
          Marco - very sorry, but I'm having trouble pinning down a precise time and date for an incident. The issue is that they weren't noted in the first place: I have emails going back a long time, but they are just from site staff with words to the effect of "Screen 10 stopped again - can you remote in and investigate?" Sometimes the message was sent several hours after the incident, or even a day later.

          I will ask the managers to try to give me as near as possible to a precise date and time if there is a repeat incident.

          That having been said, I think you're on to something with that CP950. I downloaded a log from it and uploaded that to the analyzer. It found this:


          image.png

          Both the IDF switch in the pedestal and the cable between it and the CP950 have been replaced recently. This now leads me to suspect that the network interface in the CP950 could be bad.

          The other weird thing is that the switch shows no link at all for the port that the CP950 is connected to:

          image.png

          That clearly isn't true, because I've just been using the CP950's web UI and downloading a log through it!

          I think the next step is to swap that CP950 with another screen, and see if the incidents follow it.

          Of those entries on the ARP table, 172.16.100.50 is the remote access Teamviewer PC, 172.16.100.99 is the TMS, 172.16.110.2 is the DSS220, and 172.16.102.1 is the router/DHCP server. That all looks normal to me.

          Originally posted by Caleb Williams
          Even after the change, the 200's used a hardware raid controller that exposed the array as if it were a single disk to the BIOS. That is the critical difference between the 200 and the 220. Even with later software versions, I witnessed the same issue on the 220 as Steve, where a failure of drive 1 required an entire server rebuild. In some cases even a random power bump would kill a 220.
          All that software RAID is effectively any good for is to ensure that if a drive dies during a show, that the show will not be lost as long as the server remains powered. At the next reboot, it'll fail, as others have noted. For this reason I have a complete set of three drives (in cartridges, culled from a retired DSS200), configured as screen 20. Should a manager notice a red light on one of the drive cartridges, or do their morning walk through the booth and find one of the DSS220s with a blinking GRUB prompt, I've told them to take all three drives out and replace them with the emergency set. Once it's booted, I'll go in remotely and configure them for the correct screen number, and start the content re-ingesting from the TMS. When I can get to the site, I'll then replace the bad drive, reinstall the original set, and do a from scratch software reinstall on it.

          Comment

          • Marco Giustini
            Film God
            • Jan 2020
            • 1168
            • Reading, UK

            #35
            No worries, Leo - even a date would be a starting point if you have one.

            Great that you found something - last time I had a similar issue in a cinema, turned out the network installers joined lengths of cat 6 cables by twisting the pairs by hand and adding electrical tape. The whole network had to be re-done.

            Comment

            • Caleb Williams
              Pro Film Handler
              • Jun 2022
              • 149
              • Casper, Wyoming, USA

              #36
              Originally posted by Marco Giustini
              .....network installers joined lengths of cat 6 cables by twisting the pairs by hand and adding electrical tape. The whole network had to be re-done.
              This is horrifying

              Comment

              • Marco Giustini
                Film God
                • Jan 2020
                • 1168
                • Reading, UK

                #37
                Originally posted by Caleb Williams

                This is horrifying
                Indeed. Particularly because I discovered the issue a few years after having to restart some (non cinema) servers basically every day because they kept going down! I started digging into the network and found the intel driver reporting buckets of transmission errors. Then I opened one trunk and found the mothball of electrical tape!

                I could only imagine what a Fluke would have said about that!

                Comment

                • Frank Cox
                  Film God
                  • Jan 2020
                  • 2314
                  • Melville Saskatchewan

                  #38
                  If it works, it's a Fluke!

                  Comment

                  • Harold Hallikainen
                    Film God
                    • Jan 2020
                    • 1071
                    • Tucson AZ

                    #39
                    Originally posted by Marco Giustini
                    Great that you found something - last time I had a similar issue in a cinema, turned out the network installers joined lengths of cat 6 cables by twisting the pairs by hand and adding electrical tape.
                    In our house in Arvada CO (near Denver), the lights went out in several rooms. I found in the crawl space that someone had tapped into the Romex to add lights on the back porch. The wires were just twisted and taped. I ended up adding two junction boxes since I could not pull the two pieces of original Romex close enough together to get them into a junction box with some wire nuts. So, the original "tap" was put in one junction box, then a short piece of Romex to a second junction box where the original Romex to the other rooms was joined to the short extension with more wire nuts. Not quite a gigbit network, but just twisting wires and taping them is not a great idea at any frequency.

                    Comment

                    • Leo Enticknap
                      Film God
                      • Jan 2020
                      • 3572
                      • Loma Linda, CA

                      #40
                      Originally posted by Caleb Williams
                      This is horrifying
                      I once had to do a service call to a theater in which all the stage channels suddenly went out. There was an amp rack containing a switch plus a couple of Q-Sys IOFrames that acted as a D to A for the stage channel amps, also in that rack, behind the stage. The switch was uplinked to the rack containing the core and the surround amps in the booth by a single cat6 cable running through the ceiling void of the auditorium. The theater had a rodent infestation, with the result that rats nibbled the network cable (thankfully in the area behind the screen where the damage was accessible with the ladder available.

                      For a moment I thought that I would have go the twist and electrical tape route, as I believed it to be the only solution in my tool cart. However, I then remembered that I had a box of lever nuts, (I use them frequently to extend the low voltage power cords of our IRC-28C HI/VI/CCAP emitters) and used eight of them to do a temporary repair to the cat6 cable. Communication was restored between the IOFrames and the core, audio played, and the screen was back up.

                      I advised the theater both verbally and in my service call report that once the vermin infestation was confirmed as having been eradicated, this network cable should be replaced and re-run in its entirety, ideally in a conduit (and consideration given to upgrading the system to dual redundant Q-LANs, too). I never heard anything from them since. I therefore don't know if this was ever done, or if the lever nut wire splices are still in service.
                      Last edited by Leo Enticknap; 10-08-2026, 08:14 PM.

                      Comment

                      • Marco Giustini
                        Film God
                        • Jan 2020
                        • 1168
                        • Reading, UK

                        #41
                        Originally posted by Leo Enticknap
                        I never heard anything from them since. I therefore don't know if this was ever done, or if the lever nut wire splices are still in service.
                        Having worked in cinemas for a few decades, I have a strong feeling I know the answer to that!

                        Comment

                        • Leo Enticknap
                          Film God
                          • Jan 2020
                          • 3572
                          • Loma Linda, CA

                          #42
                          Herewith an update, after a site call yesterday.

                          Firstly, an admission: when I stated earlier that I replaced all the network cables in the pedestal, I forgot to mention that "in the pedestal" excluded the cable between the IDF switch and the CP950, which is not actually in the pedestal, but a rack next to it:

                          image.png

                          At the time I skipped that one, because I didn't have a long enough patch cable or a fish tape to get it through the conduit with me, and figured that there was no way that an issue with an audio processor could affect communication between the server and the IMB (via the projector's backplane). But from Marco's latest look at the logs, it does appear that the drops in server to IMB communication happened immediately after the DSS220 tried and failed to poll the CP950 for status information.

                          Therefore, I did replace the CP950's management network cable yesterday (and might as well not have bothered with the fish tape, because the conduit is so stuffed with other cables that the RJ45 plug wouldn't go through: I had to zip tie the cable to the outside of the conduit). I also swapped SMPS cards with another projector and replaced the UPS batteries, to rule out two other dirty power possibilities, and pulled, contact cleaned, and reseated the cat1700 board in the CP950.

                          After all that was done I did a clean reinstall of the DSS220's software again, to erase all records of the CP950 communication problems and thereby ensure that any log taken after that point only records events that happened after the CP950's network cable was replaced.

                          DSS220 log downloaded immediately before the clean reinstall

                          image.png
                          The last connection loss incident it shows is on September 9 at 3.23a.

                          DSS220 log downloaded about half an hour after the clean reinstall

                          Clean so far, but that doesn't mean anything after such a short time. The acid test will be when I download another log when I'm next there (likely in about 2-3 weeks), if there are still no "loss of connection" entries.

                          CP950 log downloaded about half an hour after the DSS220 clean reinstall

                          Annoyingly the analysis doesn't give time stamps for the dropped packets, but if that warning has dropped off the analysis report when I download another after a few weeks, that would be a good sign.

                          The site staff will run some "ghost shows" with the lamp off over the next few days to see if there are any further incidents, and if there aren't any, reopen that house.

                          The one discouraging piece of evidence is that when I took the suspect network cable (between the CP950 and the IDF switch) and tested it, it appeared to be OK:

                          image.png​

                          All eight lights at both ends lit up in sequence and unison for about 2-3 minutes. Admittedly this tester is only a $20 from Amazon one - a Fluke it ain't - but I was kinda hoping to find the smoking gun with it, which would have been pretty convincing evidence that this was the problem.

                          If the incidents repeat and the times match with CP950 dropped packet errors, I'm going to suspect the CP950. It's a very early one (manufactured in February 2020 according to the label on the back), and so swapping that with another screen will probably be the best thing to try. But I'm hoping that the cable was marginal and still the cause of the trouble. Only time will tell, I guess.


                          Comment

                          • Marco Giustini
                            Film God
                            • Jan 2020
                            • 1168
                            • Reading, UK

                            #43
                            Thanks Leo,

                            Let me clarify again that my finding was honestly just a loud thinking. I noticed the 950 error logged but you were unable to give me a date and time of an incident so I couldn't confirm the 950 is in fact linked to the disconnections you've been seeing. I honestly wouldn't take that flag too seriously until we have a date/time of an incident.

                            Those errors might be "normal" - you didn't by any chance download a set of logs from another projector with an identical setup for comparison?

                            Do the loss of connection listed above coincide with a playback interruption? Did you extract the logs before wiping the DSS220?

                            On the Network tester, if that tester showed you an error, you'd have permanent lack of communication. That is a continuity test. It just tells you if you can light a lightbulb using the cable, nothing more

                            Comment

                            • Christos Gartaganis
                              Film Handler
                              • Sep 2021
                              • 53
                              • Athens Greece

                              #44
                              I am joining this topic a bit late but have you ruled out the possibility of intermittent hardware failure with the Ethernet ports, either on the Cat. 745 side or the DSS220 side?

                              Comment

                              • Marco Giustini
                                Film God
                                • Jan 2020
                                • 1168
                                • Reading, UK

                                #45
                                I believe both the 745 and the DSS have been swapped and the issue remained.

                                Comment

                                Working...