Random losses of communication between DSS220 and cat745 IMB

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Leo Enticknap
    Film God
    • Jan 2020
    • 3572
    • Loma Linda, CA

    #1

    Random losses of communication between DSS220 and cat745 IMB

    I have an ongoing problem, the first instance of which goes back to 2019, of occasional random show stops with "Transport not available - error connecting to cat745" in one screen of an 11-plex. The projector is a Barco DP2K-15C, with a Dolby DSS220 and a cat745 IMB. Sometimes the fault spontaneously clears by itself within anything from a few minutes to a few hours. If it doesn't, rebooting the server (without power cycling the projector) always clears it. The projector and server/IMB are on the current software/firmware bundles (2.10.128 for the projector and 4.9.6.4 for the DSS220).

    Nothing in either the server or projector logs sheds any light on this. Dolby have looked at the DSS220 log package (and I also put it through the online log analyzer) and can't see anything immediately before the incidents to explain them. Both I and Barco have looked at projector logs, which also have no abnormal entries at the time of the incidents.

    Between February 2019 and June this year, these incidents happened once or twice a year at most. Then, the frequency suddenly increased to the point at which the customer asked for hard core, priority troubleshooting. It's now happening roughly every week to ten days. Since June I have done the following:
    • Replaced the DSS220 server (swapped with another screen)
    • Replaced the cat745 IMB (swapped with another screen)
    • Replaced the projector's CCB (swapped with another screen)
    • Replaced the projector's backplane (new)
    • Replaced the IDF switch in the projector's pedestal (new)
    • Replaced all the Ethernet cables between the switch and the devices in the pedestal with new, and reterminated the management and media uplink cables with new RJ45 plugs.
    All the contacts of all the swapped cards were cleaned with an alcohol prep pad and then coated with DeOxit before being reinstalled.

    Unless I'm missing something, this is everything, apart from the ICP, that data traveling between the server, the IMB, and all the components of the projector that they need to be able to communicate with, touches.

    I have not tried an ICP swap so far, because all the ICPs in the complex are all well over 10 years old, which is the design life of the soldered on certificate battery that cannot be replaced. Therefore, I want them powered down as little as possible, in case those batteries are already discharged, and therefore powering down the ICP will kill it. The cat745s are in the same boat: because of the notoriously high failure rate when replacing their certificate batteries, they haven't been swapped since Dolby dropped recertification support, and the projectors left on 24/7 to preserve their lives for as long as possible. Of the 11 originally installed, 8 survive, and three have had to be replaced with new servers (an IMS2000, an IMS3000, and an ICMP+ respectively).

    Am I missing something, or is an ICP swap now the only thing left to try? Thoughts appreciated.
  • Marco Giustini
    Film God
    • Jan 2020
    • 1168
    • Reading, UK

    #2
    I'd begin with: can you ping the 745 when the issue happens? On a DSS/745 system the IPs are as below:

    DSS2xx = 130.1.1.1
    CAT745 = 130.1.1.2

    Ping from the server. If you see a reply, then I'd imagine the issue is not the network. But you say "terminated". Do you make your own cables? As much as I like doing that myself, you'd really need a Fluke or similar to make sure the cable is gigabit compliant. If not, it might work... until it doesn't. If you fiddle with linux, you might be able to confirm whether the 745 interface (the one with 130.1.1.1) shows packet errors.

    Then the 24/7 thing. I understand why that is done but... are you saying that combination is never rebooted/power cycled?

    Also, I'd probably wire the 745 directly to the server - no switches.

    Finally: did you try a full reinstall of the DSS? Granted, you replaced it so it shouldn't be the server. Yet, it's a simple-ish task to do on a DSS.

    Comment

    • Leo Enticknap
      Film God
      • Jan 2020
      • 3572
      • Loma Linda, CA

      #3
      When the issue presents, I can connect directly to the cat745 using the Dolby media block app, and the RDY light stays green.

      The video IP cable does go straight from the server to the IMB. Management and media go through the switch (separate VLANs for each).

      A full "clean install" of the DSS220 was done when I swapped them. In fact, they've all had clean installs quite frequently, because, thanks to the software RAID, that has to be done whenever a drive fails.

      I reterminated the IDF to MDF cables with new plugs and one of these. While I only tested them for continuity with a basic tester afterwards, there have been no problems with any of the other devices that use the uplink. The switches autonegotiated a gigabit full duplex connection, so I'd be surprised if that's the problem.

      The projector is power cycled as little as possible to preserve the ICP and IMB batteries: when absolutely necessary for a repair or troubleshooting, and for maintenance C and D, which require it to be powered down. Even then, I try to prepare to ensure that the power down is for as short a time as possible.

      Comment

      • Steve Guttag
        Film God
        • Jan 2020
        • 3777
        • Annapolis, MD

        #4
        I'm not seeing that on the units I support...however, most of the units I support are power cycled daily (they don't run 24/7). I've had all sort of DSS server NICs just stop working until rebooted too. Just recently, on three (out of 7) DSS200s, the Theatre NIC just stopped communicating...reboot...all is well (and on those servers, the Theatre NIC is a separate card). I'm more inclined to think it is something up with the DSS software than the hardware, besides the hardware aging.

        I have found that 4.9.5.2 seems to have fewer unexplained issues than 4.9.6.4.

        On the DSS220...you can swap drives 2 or 3 and it will rebuild. Drive 1...not so much. I just had to rebuild a RAID because drive 1 cratered.

        I don't believe the ICP is your problem or could cause what you are describing. It is the CAT745, the DSS220 MB, the connecting cable, or the software. You've changed all of the hardware and get the same results. Maybe...if you are completely unlucky, the projector is providing bad power and the CAT745 is responding to it. I JUST recently had a SMPS do that with an ICP. Barco diagnosed the problem as a bad ICP because at each failure, that board dropped out, causing the Enigma to declare no communication. However, the ICP was responding to bad power. I changed the SMPS and the problem was gone.

        So, I would suggest swapping the SMPS before swapping the ICP. I would also suggest running 4.9.5.2 over 4.9.6.4.

        Then again, swapping away from the DSS line, at this stage, is probably the better move...but to which server? I have my reservations about some of the choices.

        Comment

        • Leo Enticknap
          Film God
          • Jan 2020
          • 3572
          • Loma Linda, CA

          #5
          Thanks Steve. I didn't think of the SMPS, but it makes perfect sense. I will try that next, as it's much lower risk than the ICP (the projectors would only need to be off for a couple of minutes, and no need to disturb the ICP hardware, risk a very old and dried out solder join failing, etc.).

          I would like to see these servers and IMBs go to the e-waste bin, too (and in a Barco projector, would be inclined to recommend the ICMP, because the soldered on battery problem goes away and it's much cheaper than an ICP-D plus a third party IMS), but the customer wants to get as much life out of them as possible.

          Comment

          • Steve Guttag
            Film God
            • Jan 2020
            • 3777
            • Annapolis, MD

            #6
            The ICMP-X would make me a bit cautious only because the ICMP-XS is Barco's current/future unit...as such, I have to believe that the ICMP-X is a short-timer server. If one were to change out the projector with the ICMP-X, there are no scenarios where you would keep it. New Barco? You'd get the projector with the ICMP-XS...new non-Barco? Can't use it. At least with the IMS3000 or SR-1000, that server investment could work with ANY new projector (except the 300nit one...but anyone putting that projector in isn't thinking about keeping old stuff anyway). The IMS3000 and SR-1000 works with the existing ICP too. Have you measured the battery level? Mine seem to be doing well with most at 2.9-3V.

            Comment

            • Jim Cassedy
              Film God
              • Jan 2020
              • 1450
              • San Francisco

              #7
              I'd occasionally have this problem at one of the preview theaters I worked at which had
              DSS-220 + CAT862 connected to an NEC NC-2000C. It usually happened "overnight"
              if the system was left turned on. I'd come in the next day and get "TRANSPORT NOT
              AVAILABLE" message. I never did find the cause, but since re-booting the DSS-220
              always cleared the error, it was more of a minor annoyance than an actual 'problem'.

              Comment

              • Jason Sharp
                Film Handler
                • Apr 2022
                • 43
                • Denver, Colorado, USA

                #8
                I had a similar puzzling issue with a NEC 2000 and DSS200 CAT745 a few years back. It ended up being the female/female bulkhead network port on the projector. I removed the port and ran the ethernet cable straight to the projectors internal router. For this problem it would only cause momentary "transport not available" messages.

                With your problem lasting minutes to hours, perhaps there is a misconfigured IP address on a different device that is causing a conflict with the projector or DSS. Try giving known vacant IP addresses to the projector and DSS.

                Comment

                • Marco Giustini
                  Film God
                  • Jan 2020
                  • 1168
                  • Reading, UK

                  #9
                  Originally posted by Leo Enticknap
                  When the issue presents, I can connect directly to the cat745 using the Dolby media block app, and the RDY light stays green.


                  I reterminated the IDF to MDF cables with new plugs and one of these. While I only tested them for continuity with a basic tester afterwards, there have been no problems with any of the other devices that use the uplink. The switches autonegotiated a gigabit full duplex connection, so I'd be surprised if that's the problem.
                  Your tests says the 745 is still responding but a network ping would dismiss (partly) a network issue.

                  I say partly because I know you never had issues with your cables but when it comes to Gigabit, it doesn't take much to upset the signal. A noisy ethernet cable can be dealt with by the network interface until it cannot. It's unlikely but I'd move to a pre-made, tested cable for this test.

                  I started crimping my own cables many years ago for work - then I realised that without a proper tool, it was just too risky. Any random issue and the network would be immediately be the number one suspect.

                  Finally, as Steve says, those servers were not bullet proof and I'd say a weekly reboot was (and still is) strongly recommended to clear "errors"

                  Jim,
                  It's a bit foggy but I believe the DSS does something overnight - as a start it would renew the TLS keys which would result in a momentary loss of connection. And we know that momentary in IT might mean permanent

                  Comment

                  • Leo Enticknap
                    Film God
                    • Jan 2020
                    • 3572
                    • Loma Linda, CA

                    #10
                    Originally posted by Steve Guttag
                    The ICMP-X would make me a bit cautious only because the ICMP-XS is Barco's current/future unit...as such, I have to believe that the ICMP-X is a short-timer server.
                    But the XS won't work in a Series 2 projector. The IMS3000 and SR-1000 have also been in production for the best part of a decade now, too (nine years in the case of the IMS3000: the first one I installed was on June 12, 2017 - ironically, it survives to this day, still running on its original boot flash drive!), so it can't be long before Dolby and GDC launch replacements.

                    Originally posted by Jim Cassedy
                    It usually happened "overnight" if the system was left turned on. I'd come in the next day and get "TRANSPORT NOT AVAILABLE" message. I never did find the cause, but since re-booting the DSS-220 always cleared the error, it was more of a minor annoyance than an actual 'problem'.
                    In my case it's causing shows to stop midway, meaning that we have to get to the bottom of it.

                    Originally posted by Jason Sharp
                    With your problem lasting minutes to hours, perhaps there is a misconfigured IP address on a different device that is causing a conflict with the projector or DSS. Try giving known vacant IP addresses to the projector and DSS.
                    I did wonder if there was an address conflict, too. Sorry - forgot to mention in my original post. After one incident, we disconnected both the management and media LAN uplinks to the MDF switches, and it happened again a couple of days later. I am 100% positive that there is nothing else with the same address as the projector or server connected directly to the IDF switch, and so I think an IP address conflict can be ruled out.

                    Originally posted by Marco Giustini
                    A noisy ethernet cable can be dealt with by the network interface until it cannot. It's unlikely but I'd move to a pre-made, tested cable for this test.
                    Not an option, because the IDF to MDF cable runs from that pedestal are around 200 feet through a ceiling void and over the top of the lobby between the two booths. They simply land in the pedestal with a plug: the original installers did not terminate them into a patch panel. Replacing the entire cable run would mean getting contractors in. In any case, when we pulled them for a couple of days, there was a repeat incident while the uplinks were disconnected. Also, neither the management nor the media uplinks are showing any bad packets at the IDF switch:

                    image.png

                    Good idea on a weekly reboot for those servers. I believe that the GDC TMS can schedule and automate that, and will look into it - thanks.

                    In the meantime, the next time I can get to the site, I'm going to try an SMPS swap if the manager is OK with the risk that the random show stops could move to the screen that receives the suspect SMPS.

                    Comment

                    • Marco Giustini
                      Film God
                      • Jan 2020
                      • 1168
                      • Reading, UK

                      #11
                      Originally posted by Leo Enticknap
                      Not an option, because the IDF to MDF cable runs from that pedestal are around 200 feet through a ceiling void and over the top of the lobby between the two booths. They simply land in the pedestal with a plug: the original installers did not terminate them into a patch panel. Replacing the entire cable run would mean getting contractors in. In any case, when we pulled them for a couple of days, there was a repeat incident while the uplinks were disconnected. Also, neither the management nor the media uplinks are showing any bad packets at the IDF switch:
                      Gotcha. On that relatively long run, I feel that having it tested with a proper network tester is a must. In the past, on sites I worked, a team from the networking company would crimp all the cables, another would test and certify each line. I do remember a site where the certification failed on basically everything and the whole network had to be re-terminated/crimped properly.

                      If you don't want to power cycle the units, I'd imagine you could at least reboot them via software. The DSS will accept a reboot command, the projector should have a reboot command as well. The 745, I don't know.

                      If the shows fails mid-show, I'm surprised the logs don't give a clue on what happens?

                      Comment

                      • Leo Enticknap
                        Film God
                        • Jan 2020
                        • 3572
                        • Loma Linda, CA

                        #12
                        It looks like I can automate a weekly reboot via the GDC TMS using its "auto PMA" feature, but it'll be a bit of a science project, requiring a Python script to be written. I'm going to look at this further as and when time allows.

                        I'm as surprised as you are as to the lack of leads from the logs. The disconnection events generate an entry recording that they happened, but with no clue as to what provoked it. Here is an analysis of a log taken after an incident on June 28 (the last time I bothered to download a log, as you can't remotely from a DSS, and if this one wasn't going to tell us anything, the later ones likely won't, either). Here is an actual, raw entry:

                        image.png

                        It was the simple recording of a disconnection that led me to suspect a networking problem, hence replacing the switch and cables and reterminating the uplinks; and when that didn't stop the trouble, swapping both the DSS220 and the projector's CCB. There is also an entry for "Device[audio]" disconnecting, which I presume means the CP950, and which I thought strengthened my case for there being a networking problem. But all the LAN components between the DSS220, the projector, and the cat745 have been replaced with new, and IP address conflicts ruled out; yet the incidents persist.
                        Last edited by Leo Enticknap; 09-26-2026, 09:44 AM.

                        Comment

                        • Steve Guttag
                          Film God
                          • Jan 2020
                          • 3777
                          • Annapolis, MD

                          #13
                          I'm not sure a soft-reboot will be as good is a proper power cycle where the memory and everything lose power. So, if there is a leak, that will be addressed as well. Again, my system power cycle weekly (the servers) and the projectors power cycle daily with the open times of the theatre.

                          Comment

                          • Tony Bandiera Jr
                            Expert Film Handler
                            • Jan 2020
                            • 720
                            • Moreland, Idaho

                            #14
                            Originally posted by Leo Enticknap
                            I have an ongoing problem, the first instance of which goes back to 2019, of occasional random show stops with "Transport not available - error connecting to cat745" in one screen of an 11-plex. The projector is a Barco DP2K-15C, with a Dolby DSS220 and a cat745 IMB. Sometimes the fault spontaneously clears by itself within anything from a few minutes to a few hours. If it doesn't, rebooting the server (without power cycling the projector) always clears it. The projector and server/IMB are on the current software/firmware bundles (2.10.128 for the projector and 4.9.6.4 for the DSS220).

                            Nothing in either the server or projector logs sheds any light on this. Dolby have looked at the DSS220 log package (and I also put it through the online log analyzer) and can't see anything immediately before the incidents to explain them. Both I and Barco have looked at projector logs, which also have no abnormal entries at the time of the incidents.

                            Between February 2019 and June this year, these incidents happened once or twice a year at most. Then, the frequency suddenly increased to the point at which the customer asked for hard core, priority troubleshooting. It's now happening roughly every week to ten days. Since June I have done the following:
                            • Replaced the DSS220 server (swapped with another screen)
                            • Replaced the cat745 IMB (swapped with another screen)
                            • Replaced the projector's CCB (swapped with another screen)
                            • Replaced the projector's backplane (new)
                            • Replaced the IDF switch in the projector's pedestal (new)
                            • Replaced all the Ethernet cables between the switch and the devices in the pedestal with new, and reterminated the management and media uplink cables with new RJ45 plugs.
                            All the contacts of all the swapped cards were cleaned with an alcohol prep pad and then coated with DeOxit before being reinstalled.

                            Unless I'm missing something, this is everything, apart from the ICP, that data traveling between the server, the IMB, and all the components of the projector that they need to be able to communicate with, touches.

                            I have not tried an ICP swap so far, because all the ICPs in the complex are all well over 10 years old, which is the design life of the soldered on certificate battery that cannot be replaced. Therefore, I want them powered down as little as possible, in case those batteries are already discharged, and therefore powering down the ICP will kill it. The cat745s are in the same boat: because of the notoriously high failure rate when replacing their certificate batteries, they haven't been swapped since Dolby dropped recertification support, and the projectors left on 24/7 to preserve their lives for as long as possible. Of the 11 originally installed, 8 survive, and three have had to be replaced with new servers (an IMS2000, an IMS3000, and an ICMP+ respectively).

                            Am I missing something, or is an ICP swap now the only thing left to try? Thoughts appreciated.
                            Did you replace the little internal jumper cable in the DSS200? It went from the back panel to the Cat 682 (Does that same connection get used with a cat 745?, i.e. linking from the motherboard to a back panel jack? Put it another way, is there an internal jumper cable betwixt the MB and the back panel connecting the cat745?) that caused this exact error on EVERY DSS200 I installed back in the day, replacing that internal cable with a higher quality one stopped the issue.

                            Other than that, as you know my last DSS200 was at a certain screening room way back in the day.

                            Comment

                            • Leo Enticknap
                              Film God
                              • Jan 2020
                              • 3572
                              • Loma Linda, CA

                              #15
                              Originally posted by Steve Guttag
                              I'm not sure a soft-reboot will be as good is a proper power cycle where the memory and everything lose power. So, if there is a leak, that will be addressed as well. Again, my system power cycle weekly (the servers) and the projectors power cycle daily with the open times of the theatre.
                              The pedestals as they are currently configured do not allow for this. To do that would require new hardware, and because the projector is 240V, something that would enable it to be power cycled on a schedule would not be cheap. A relatively inexpensive WattBox would enable the 120V stuff (server and CP950) to be power cycled automatically on a schedule, though. As this fault only affects one of the 11 screens, I don't think that the theater manager would be keen to spend significant money on an attempted fix that may or may not work. I already persuaded him to do that when we swapped the backplane (although that having been said, it wasn't money wasted, because the life expectancy of a factory original backplane in 14-year old Series 2 Barco is similar to that of Ted Bundy shortly after being strapped into the electric chair).

                              The more I think about this, the more I believe that an SMPS swap, as you suggested, is the one remaining fix to try that is clearly a possibility and would not require money to be spent. I'm going to try to persuade them to let me try it when I'm next there (likely back end of this coming week or early the week after).

                              Originally posted by Tony Bandeira Jr.
                              Did you replace the little internal jumper cable in the DSS200?
                              I wish there was one to replace! Sadly, this is a DSS220, not a 200. The 220 is a "200 lite," minus the hardware RAID card and cat862 media block. It's basically just a server motherboard, power supply, and three-drive RAID cage, and will only work with a cat745 (the 200 gave you the option of either using the internal cat862 media block via that little jumper, or an external cat745 IMB). It was you or someone else on F-T that flagged up that internal jumper as the cause of multiple cat862 disconnections: after reading that, I replaced it on any DSS200 that I had to open up for any reason. At the factory, they stupidly zip tied it in such a tight turn that it was effectively a fold. No wonder so many of them failed.

                              Comment

                              Working...