Job processing delays with possible failures
Updates
The data flow pipelines are fully operational since 12hrs ago. We kept monitoring the platform and fixing actions which got unexpectedly broken during that time.
We fixed the root configuration issue which caused this incident and which revealed also a set of subsequent bugs, which we gradually fixed too. Now the platform is again fully operational and our team restarted the failed data pipelines which failed during the incident time. We will keep continuously observing the platform health and update the status if needed.
We're currently observing significant amount of action failing and being delayed in processing. The root cause is invalid configuration of our apps responsible for processing the data and their scaling configuration. Our engineers already work on solving the issue. The queued data flows started to be processed again after they were paused. We expect some actions may fail due to previously unseen errors which is expected. Such actions will be re-started by dataddo team after the queues is solved automatically. No need to do anything on users side. Feel free to contact us in case of any questions.