Following up on the major Xbox outage from a couple of weeks ago, the CTO of the division, Scott Van Vliet, has now offered an update. In a lengthy social media post, Vliet once more confirmed that the outage was caused by an “underlying service” that is used by Xbox to validate purchased and Game Pass content, as well as some services on Xbox for PC.
Another bug was discovered in the Xbox client software stack that further compounded the issue, where services in an already unresponsive state also ended up causing other license checks to fail. This, in turn, led to a cascading failure of services, affecting users’ ability to even launch games that they already owned.
Since then, Vliet has noted that Xbox has transferred ownership of the licensing services that caused the issues to its in-house team, and engineers have been working on making the software stack more resilient to these kinds of failures. The bug that was discovered in the stack has also been fixed with an update that has started rolling out on August 10th.
Looking to the future, Vliet has explained that the engineering and operations teams at Xbox are actively working to make sure that the Xbox stack at large is more reliable. More updates about how this will be achieved are slated to be revealed in the coming weeks and months. Updates fixing these issues will also be accompanied by new features and quality-of-life improvements.
The Xbox service failure was a major issue in July, with online services being down for around 15 hours. During this time, players discovered that, along with not being able to access network features, they also weren’t able to launch digital copies of games that they had bought. At the time, Vliet had called the outage “unacceptable”, and said that the team was committed to ensuring that it doesn’t happen again.
While Vliet acknowledged that the service outages on their own were bad, the fact that they were able to cause a larger software stack collapse due to cascading failure made it worse. Along with this, he also said that the company wants to be more transparent with issues like this going forward.
“I care less about the one-line root cause and more about the real questions: why a failure in one service was able to take down this much, why recovery took as long as it did, and what we change so a single point of failure can’t ruin your night again. That means hardening the dependencies underneath sign-in and game launch, improving how we detect and contain this class of failure, and being faster and clearer with you when something breaks.”
The outage happened right in the middle of larger conversations taking place all over the industry about the pros and cons of digital media versus physical. With Sony announcing that it is going to end the production of PlayStation discs, many users have been expressing concern over how this would affect the ownership of games, and whether services would be able to stand up to system failures.















