Random losses of communication between DSS220 and cat745 IMB

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Marco Giustini
    Film God
    • Jan 2020
    • 1168
    • Reading, UK

    #16
    The logs says "LeaseRenewTask". As far as I know that is the TLS handshaking renewing every 24 hours. That said, is that supposed to happen with a 745? I thought it was an 862 thing. Or maybe it's referring to the IP lease - which I am not sure as the 745 IP should be fixed. Might be a red herring.

    I would look into the 745 logs - it's a separate file - and see what happens at the same time, bearing in mind they might be some time off, it's always been a pain to match logs as the server might show logs to your current time zone but most internal logs will be UTC.

    Logs are going down to up. I see that before the projector, your sound device is losing communication as well? If that is the case, that would be a huge clue as I'd imagine the sound device is nowhere near the 745 network, right?

    Comment

    • Leo Enticknap
      Film God
      • Jan 2020
      • 3572
      • Loma Linda, CA

      #17
      I don't know if it's TLS or DHCP, but I'd be surprised if it's DHCP. That entry refers to a 192.168... address. In that theater, the management/auditorium LAN is 172.16.1XX.Y / 16, where XX is the auditorium number (so, for example, Screen 8's projector is 172.16.108.3). The media/theater LAN follows the Dolby DSS default of 192.168.241.X, where X is 2 for the TMS and the screens start at 3. This is because this place originally had a Dolby DSL200 TMS, which was later replaced with GDC. I saw no reason to re-IP the media/theater LAN, and so I just left it as is.

      The theater/media LAN is static: there is no DHCP server on it, anywhere. However, the DSS servers have that address setting set to default. They always set themselves to the what they should be during bootup (for example, 192.168.241.10 for Screen 8), but I wonder if under the old DSS/DSL ecosystem, the DSL acted as a DHCP server that gave each screen server its address, and that it would only give itself an address based on the configured screen number if its attempt to pull a DHCP address timed out?

      I suppose I could try setting those addresses statically in the server, but this does not explain why there are six other identical screens in terms of equipment configuration (DP2K-XXC, DSS220/cat745 and CP950), and this fault has not presented in any of them.

      In terms of the sound device, there was a bug affecting DSS servers connected to CP950s, that was fixed in CP950 version 2.3.2.5:

      image.png

      ...but this CP950 is on 2.3.3.7, as are all the others in the complex. Furthermore, this is not an accurate description of the fault that is actually happening (complete cessation of playback, not just audio). And again, the problem isn't happening in any of the other DSS + CP950 screens.

      Comment

      • Marco Giustini
        Film God
        • Jan 2020
        • 1168
        • Reading, UK

        #18
        With TLS I meant Transport Layer Security - the encryption for the media block. Now, that was maybe a thing of the past, I am not sure how the 745 behaves on that. With the 862, the TLS encryption key would be renewed every 24 hours - which was usually done at boot or in the middle of the night. I believe there were occurrences where the media block would stop working while the DSS renewed the TLS key. It was addressed via software at some point - but it's all very foggy! Steve will know for sure

        I'd imagine 192.168.241.12 is the server itself? It's how the log calls the server I guess.

        I'm not sure I made my point though. From the very limited section of logs you showed, the sound processor AND the projector BOTH are tagged as disconnected by the server. In fact, the audio device is disconnected 30 seconds before the projector is also marked as disconnected.

        Without knowing much about this issue and your network, that seems to suggest a more general network issue rather than a 745 issue.

        It's also compatible with "we swapped server and 745 and the issue remained"

        I hope you didn't say that already: can you quickly go through how the server, 745, projector, sound processor and TMS are connected in this auditorium from a network perspective?

        Comment

        • Steve Guttag
          Film God
          • Jan 2020
          • 3777
          • Annapolis, MD

          #19
          I want to say that all Dolby DSS servers re-establish TLS security at 4am and 7am, daily and I don't think the CAT745 has any bearing on it. Additionally, all servers are supposed to re-establish TLS security within 6-hours of a show start (normally satisfied when the show ends and a new show is loaded). Marathon shows can get burned by the 6-hour limit if three shows are built up as a single long performance.

          Comment

          • Marco Giustini
            Film God
            • Jan 2020
            • 1168
            • Reading, UK

            #20
            Thanks Steve,

            I'm not sure the LeaseRenewTask is the TLS one - even if it was, I think it was a byproduct of the disconnection. Projector is disconnected -> TLS security to be re-established.

            Leo, feel free to post the logs here if you are comfortable sharing them. I am no expert but I like digging into logs sometimes.

            Comment

            • Leo Enticknap
              Film God
              • Jan 2020
              • 3572
              • Loma Linda, CA

              #21
              That would suggest that the lease renewal is DHCP, unless there is a third type that I'm not aware of.

              Originally posted by Marco Giustini
              It's also compatible with "we swapped server and 745 and the issue remained. I hope you didn't say that already..."
              I did and it did - sorry!

              Originally posted by Marco Giustini
              Can you quickly go through how the server, 745, projector, sound processor and TMS are connected in this auditorium from a network perspective?
              The affected auditorium is 10. Management (auditorium in DSS terminology) IDF switch = 172.16.110.1 ( / 16); DSS220 = .2, projector = .3, CP950 = .4 / IRC-28C CCAP transmitter = .5 / amplifiers (LEA, 2 x CX704s and 1 x CX354) = .6 through 8. This is uplinked to the MDF management switch.

              Media (theatre in DSS terminology) = 192.168.241.1 (MDF switch) ; .2 (TMS) .3 (Screen 1's DSS220), etc. up to 13 (Screen 11's server).

              The automation controllers are all connected to the DSS220s via RS232 - they are MiT IMC-2Bs, which do not have Ethernet connectivity and are serial only. In screens that have non-DSS servers without RS232 interfaces, the automations are connected through these guys, which have the address .9 on management.

              The TMS communicates with the screen servers through the media/theatre LAN (these addresses are specified for the server entries), and with the projectors and audio processors through the management/auditorium LAN.

              Originally posted by Marco Giustini
              Leo, feel free to post the logs here if you are comfortable sharing them.
              The one quoted earlier is here - it's too big to attach to this post. You can download the actual package by clicking "Download log package," top right. If you do manage to find a smoking gun, you will be one up on Dolby!

              I'm going to be back at this site on October 9, and have the manager's OK to try a SMPS swap.

              Comment

              • Marco Giustini
                Film God
                • Jan 2020
                • 1168
                • Reading, UK

                #22
                Originally posted by Leo Enticknap
                I did and it did - sorry!
                I know, I was referring to sharing a precise description of your network

                The affected auditorium is 10. Management (auditorium in DSS terminology) IDF switch = 172.16.110.1 ( / 16); DSS220 = .2, projector = .3, CP950 = .4 / IRC-28C CCAP transmitter = .5 / amplifiers (LEA, 2 x CX704s and 1 x CX354) = .6 through 8. This is uplinked to the MDF management switch.

                Media (theatre in DSS terminology) = 192.168.241.1 (MDF switch) ; .2 (TMS) .3 (Screen 1's DSS220), etc. up to 13 (Screen 11's server).
                So Auditorium network 172.16.110.10 is going into a switch where all the auditorium stuff is connected. The same switch is also linked to another Main switch which also has the Theatre network connected to it.

                Theatre network is 192.168.241.1 and goes into the aforementioned switch.

                Cat745 is wired directly to the server via uninterrupted cable.

                Is the above correct?

                If it is, I would temporarily sever the link between the IDF switch and the MDF switch. You can still access the projector, sound processor, 950 etc via IP forwarding which the DSS220 has enabled by default. You will have to set a Static Route on whatever needs to access those devices.

                Basically it will instruct them to access them VIA the server.

                If that is doable, it would keep the Auditorium network nice and isolated.
                Is the MDF switch configured to have VLANs to keep Auditorium and Theatre networks separated?

                I know: "but the other screens are doing fine with the same config". Yet...

                The one quoted earlier is here - it's too big to attach to this post. You can download the actual package by clicking "Download log package," top right. If you do manage to find a smoking gun, you will be one up on Dolby!
                Got it thanks.

                I would need to know one more thing:

                - Date and time of one occurrence. is June 28th 1.12PM one of them I could use for digging? The more you could tell me about one of those occurrences (the show stopped, the server was rebooted 3 times etc), the easier is going to be to dig into the logs.
                - The time the show which was interrupted started
                - The name of the show/content that was being played. It would help untangling the events

                Thanks!

                Edit:

                I see the server refreshes its cues every hour or so. When it connected to the 950 to get its status, the 950 systematically fails to connect, the server throws an error but then recovers and manages to connect and get the macro names. This happens every hour.

                On June 28th at 1.12pm, the server starts the refresh cues task, the 950 fails to connect as usual but then everything else disconnects.
                HOWEVER, I don't think anything was playing at that time. So I'll away for your feedback on a date and time (plus other details) when the issue happened.
                Also - because the 950 wasn't a thing when I used to deal with DSS' - do you have a log from another, working auditorium for comparison? One with a 950 as well.

                Of course all those "errors" might be "expected", who knows.

                Code:
                2026-06-28 12:49:54,465 INFO  [Timer-3] streaming.StreamManager - refreshCues
                2026-06-28 12:49:54,465 INFO  [Timer-3] automation - Connected to device cp950:/172.16.110.4:0
                2026-06-28 13:11:00,393 ERROR [Execute getDeviceState Thread] dolby.AbstractCP950Device$GetDeviceState$1 - Error getting device state
                Last edited by Marco Giustini; 09-29-2026, 03:14 AM.

                Comment

                • Leo Enticknap
                  Film God
                  • Jan 2020
                  • 3572
                  • Loma Linda, CA

                  #23
                  Many thanks for going through that log package.

                  Originally posted by Marco Giustini
                  So Auditorium network 172.16.110.10 is going into a switch where all the auditorium stuff is connected. The same switch is also linked to another Main switch which also has the Theatre network connected to it.
                  There is no physical connection between the auditorium (management in DCI speak) and theater (media) LANs, other than that the TMS, remote access PC, and screen servers are connected to both through separate NICs in a single device. But there is no connection at the MDF switches.

                  Originally posted by Marco Giustini
                  I would temporarily sever the link between the IDF switch and the MDF switch.
                  That was done in an earlier stage of the troubleshooting, and there was an incident while the uplink was disconnected.

                  Originally posted by Marco Giustini
                  I see the server refreshes its cues every hour or so. When it connected to the 950 to get its status, the 950 systematically fails to connect, the server throws an error but then recovers and manages to connect and get the macro names. This happens every hour.
                  So this could be CP950-related? That would explain a lot, but not why several other screens in the complex also have a DSS220 with a CP950, but no random show stops. The next step, I guess, is to compare the configuration settings in 10's CP950 with those in a screen that has not had any trouble, and see if I can find any mismatch. Thanks again - will do when I get a moment.

                  Comment

                  • Marco Giustini
                    Film God
                    • Jan 2020
                    • 1168
                    • Reading, UK

                    #24
                    Last time I played with DSS logs, the 950 wasn't a thing so my experience on this is limited. If not mistaken, the 950 was added when the DSS was already unsupported so I wouldn't be surprised if some minor bugs were left in the process and those errors might be red herrings.
                    Also, logs (and DSS logs in particular) are notorious for showing "normal errors"

                    I wouldn't focus too much on my finding yet, I am still waiting for a precise date and time of one of those incidents to try to find out more.
                    And, if you can, a set of logs of a "working" auditorium.

                    Comment

                    • Bruce Cloutier
                      Pro Film Handler
                      • Jan 2020
                      • 429
                      • Pittsburgh, PA USA

                      #25
                      Without reading all of this, I thought to mention that we just worked a Barco connection issue.

                      A summary: When our connection to Barco is quiet and normal Keep_Alive ACK exchanges were ongoing, the Barco stops responding to the Keep-Alive. After several unanswered ACKs the JNIOR decides that the connection has gone away. We generate the obligatory exception. The JNIOR attempts to reconnect. The Barco then refuses to cooperate. We think it is because it thinks that there is still an active connection to us.

                      The issue/ticket was posted here this morning although I know that the guys have been investigating this for a while. We decided that when JANOS goes to drop the connection that it believes has gone away, we will have it fire out an RST packet. That just in case the Barco is still there.

                      Otherwise, the guys will be having more conversations with Barco.

                      I don't know what network stack Barco is relying on but something isn't working reliably there.

                      I don't know if this relates to any of your issues above. Just adding to the conversation.

                      We have a Release Candidate build of JANOS that could be tested. That is up to Kevin and his team. We are going to release JANOS v2.6 by the end of October.

                      Comment

                      • Leo Enticknap
                        Film God
                        • Jan 2020
                        • 3572
                        • Loma Linda, CA

                        #26
                        Originally posted by Marco Giustini
                        I wouldn't focus too much on my finding yet, I am still waiting for a precise date and time of one of those incidents to try to find out more. And, if you can, a set of logs of a "working" auditorium.
                        That is going to take a while. I won't be able to get to the site until the end of next week, and logs cannot be downloaded out of a DSS server remotely: the only way to obtain one that I know of is by shoving a USB stick up its bum and then pushing a button on the VNC UI that generates the log package onto it. At least, that is the only published method: the old Dolby/Doremi TMS was able to download logs out of a DSS over the network, but the way it did that is not documented anywhere that I have been able to find, and the current GDC TMS at that site cannot do it. So I don't have a log package from the affected server taken more recently than June as of now.

                        Originally posted by Marco Giustini
                        If not mistaken, the 950 was added when the DSS was already unsupported so I wouldn't be surprised if some minor bugs were left in the process and those errors might be red herrings.
                        That is my recollection, too: the final two software releases for the DSS line came sometime after official discontinuation, and all they did was to add support for the CP950 (4.9.5.2) and 950A (4.9.6.4).

                        Originally posted by Bruce Cloutier
                        Without reading all of this, I thought to mention that we just worked a Barco connection issue.
                        Interesting: I've had two similar issues recently. In the first, a Series 4 Barco controlled by an RTI integration/front end system was failing, in about 50% of instances, to apply the PCF file when changing from an HDMI macro to a DCI/DCP macro (with the result that the color space was wrong when a DCP started to play). I think it was a timing issue, because it was eventually fixed by moving the PCF select command to the last entry in the macro. In the second, a Series 2 Barco in a residence theater controlled by a Savant integration/front end system sometimes fails to receive the lamp off command, with the result that the lamp stays on, sometimes for days on end. The projector's log shows that no lamp off command is ever received. The company that looks after the Savant system is currently investigating on their end.
                        Last edited by Leo Enticknap; 09-30-2026, 12:15 PM.

                        Comment

                        • Marco Giustini
                          Film God
                          • Jan 2020
                          • 1168
                          • Reading, UK

                          #27
                          Originally posted by Leo Enticknap

                          That is going to take a while. I won't be able to get to the site until the end of next week, and logs cannot be downloaded out of a DSS server remotely: the only way to obtain one that I know of is by shoving a USB stick up its bum and then pushing a button on the VNC UI that generates the log package onto it. At least, that is the only published method: the old Dolby/Doremi TMS was able to download logs out of a DSS over the network, but the way it did that is not documented anywhere that I have been able to find, and the current GDC TMS at that site cannot do it. So I don't have a log package from the affected server taken more recently than June as of now.
                          I cannot remember to be fair whether there was a way to save logs remotely.

                          Edit: Jupiter Client. If you log into the server using Jupiter Client remotely, logs will be downloaded on your PC.

                          June is fine, I don't need new ones. I just need to know when in June the issue happened
                          Last edited by Marco Giustini; 09-30-2026, 01:14 PM.

                          Comment

                          • Caleb Williams
                            Pro Film Handler
                            • Jun 2022
                            • 149
                            • Casper, Wyoming, USA

                            #28
                            After reading all of this I keep coming back to the physical connection. If you can log in to the server via ssh, or just use one of the unused tty's (get to it using the keystroke cntl+alt+f<some-number>) and start a ping to the cat's address. If the ping fails around the same time then you know the link is being dropped and that is what's causing your transport error. When we ran Cat745's with DSS220's we quickly learned to replace the DSS220's with DSS200's with the Cat862 removed because the DSS200 platform was more reliable. DSS220's had a lot of issues with disk corruption and network interfaces randomly going down. (Especially during brown-outs because of the software raid. Instead of a dedicated IDE bootloader like the DSS200, they only write the bootloader on the first drive. If it dies or gets corrupted, hello server rebuild).

                            I strongly encourage the at minimum weekly reboot for the DSS220. It's a little hacky because of the union-fs but you can create a crontab that will do the reboot automatically for you directly on the server if there's not another option.

                            I'm not sure how the manager would feel about this, but if it was me I would install a temporary drop box (think nuc or raspberry pi) attached to the theater network and the internet to give you remote access, that way you can get in and see what it's doing right when the error happens instead of trying to poreing through hours of logs. I, like Marco, enjoy looking at logs, but the DSS variants were exceptionally verbose and thus contain a lot of useless information.

                            Comment

                            • Marco Giustini
                              Film God
                              • Jan 2020
                              • 1168
                              • Reading, UK

                              #29
                              The DSS200 stopped using the dedicated IDE drive at some point in the software development. Everything was on RAID.
                              I wasn't aware that losing disk 1 of the 220 would cause a catastrophic failure - are you sure about that, it would make the system very very fragile!

                              Yes, DSS logs are very chatty - hence I like them!

                              Comment

                              • Caleb Williams
                                Pro Film Handler
                                • Jun 2022
                                • 149
                                • Casper, Wyoming, USA

                                #30
                                Originally posted by Marco Giustini
                                The DSS200 stopped using the dedicated IDE drive at some point in the software development. Everything was on RAID.
                                I wasn't aware that losing disk 1 of the 220 would cause a catastrophic failure - are you sure about that, it would make the system very very fragile!
                                Even after the change, the 200's used a hardware raid controller that exposed the array as if it were a single disk to the BIOS. That is the critical difference between the 200 and the 220. Even with later software versions, I witnessed the same issue on the 220 as Steve, where a failure of drive 1 required an entire server rebuild. In some cases even a random power bump would kill a 220. In essense, the boot process of the 220 was something like

                                PWR-ON
                                |
                                BIOS
                                |
                                Disk 1 - Booloader
                                |
                                Kernel & init scripts <-- software raid isn't activated until this point
                                |
                                Dolby app stack
                                |
                                System booted

                                I remember adapting the build script because we were trying to make a DSS200 think it was a show library without changing the raid controller, and being a little confused that they didn't write the bootloader to all three disks on a 220. Perhaps it was just an oversight that was never corrected. I also remember reaching out to support about some bugs in that build script and they did not seem interested in any fixes.

                                In the long run, we found that the DSS200+cat745 was much more resilient to random power bumps and drive failures. It was providence, because around the same time we had several cat862's die. The migration just made sense.


                                None of that helps Leo unfortunately. Hopefully it's something simple like the tab on a network termination is unlatched or broken.

                                Comment

                                Working...