Guest View: Preparing for the online doorbusters

Column

Published: November 13th, 2015

With Halloween over, an even scarier event is on the horizon for IT staff: the online doorbusters that come with the winter holiday season. The term “doorbuster” applies (sometimes literally) to brick-and-mortar stores during the holiday season. It also translates well to the information technology systems that get stressed heavily on big online sales days such as Black Friday and Black Monday.

Preparing for this year: Battle School
Time is short, and most likely your software and IT systems are as well prepared as they’re going to be for the coming storm. So what can you do in the remaining time to better prepare for this year’s doorbusters?

In IT terms, doorbusting could be any manner of system failures that may occur; for example, as system transaction levels reach new peaks. Many failures are recoverable, but only if your team is ready and reacts well in high-stress situations. (Because unrecoverable failures are by definition beyond repair, I won’t address them in this article.)

A good place to focus effort with just weeks to go is training and human preparedness. The Battle School style of training is a good way to increase your team’s preparedness quickly.

What is the Battle School style of training?

It is scenario-based. Architects and other IT staff dream up a variety of potential failure scenarios for your environment. For each scenario, turn it into a discussion exercise by documenting:

What indications operators would see when it breaks
What things an operator look should for to gather data about what’s really broken
How the operator should gather more information: Should he or she consult another person, or is there a quick way to get answers from a computer? What logs can be checked, and where are they?
What corrective actions should be contemplated or executed? Who should do them, and how?

It is pull training. Pull training seeks to pull information from the brains of sharp folks rather than getting them to produce it. Typically software and IT organizations have staff at mixed levels of confidence and skill in your systems. Pull training helps spread the wealth of knowledge from the hot shots to others.

The scenarios you plan put your hot shots in a situation where they respond to “What if?” scenarios, which turns out to be a great way to extract knowledge from technical workers—especially hard-won critical-thinking skills and technical details that might matter during a system outage or other calamity. It’s much easier than getting them to write everything they know on a wiki or series of e-mails.

It’s a controlled simulation experience. When you run a Battle School training event, the participants don’t know much about the scenarios at the start. The event consists of the set of scenarios you have dreamed up and a few hours of training content. Start each exercise by discussing only what the operator indications would be for the given problem. That is, describe what participants would see, not what actually happened (per the scenario). This style helps the team gain the critical skills and learn the technical details to help determine what broke in your environment, and then learn how to fix it. Limit your scenarios to those that are recoverable.

It has collateral benefit. The work to develop and talk through scenarios leads to training events that prepare your team to handle such challenges as system failures or outages. There’s also a big collateral benefit; while coming up with the scenarios and preparing the discussion exercises, your team will uncover weaknesses in your systems. In some cases, hot shots might opt to find ways to prevent problems, unless prohibited by moratoriums on system changes as the online doorbusters approach.

Hands on?
A Battle School training event can be put together in short order with discussion-only exercises. There’s a lot of value in that step alone. Putting together a more realistic, hands-on simulation is another matter entirely. That takes more extensive preparation but may well be worth it. Consider whether a hands-on event makes sense in your environment.

Modern cloud infrastructures are changing and improving constantly. But end-user usage patterns are finding new ways to stress the systems.

The holiday season often stresses IT systems up to and beyond their capacity. Realistic stress-testing is both expensive and elusive, especially for modern cloud-based hardware topologies with shared-everything infrastructure. But it is essential for businesses that rely on the holiday season and its higher spending levels to help their cash flow.

Cloud infrastructure has come a long way in capacity, performance and reliability. But generally cloud solutions tend to be more cost-optimized than performance-optimized. Figure out what aspects of your technology stack, at least those that you have control over, are going to need operational care when sales orders or other transactions come piling in. Many organizations these days have systems that are “half” in the cloud. If that’s you, consider whether that transitional state may actually translate into service recovery options if cloud infrastructures have challenges this year.

Preparing for next year: Emergency bug-fixing strategy
What about longer-term plans for next year and beyond?

First, if you’re not already on the agile and DevOps bandwagon, get on it. Agile development promotes releasing code quickly and incrementally, aiming for a larger number of smaller, lower-impact and lower-complexity releases. Many organizations impose moratoriums on system and software changes as they near critical times, such as Black Friday, when loads are extreme. This tactic helps reduce the guesswork for IT operational staff in figuring things out when they break.

But such moratoriums typically have exceptions for a true emergency. You might encounter a bug that only uncovers itself in the systems load of the holiday season. Your superhero coders have a better chance at delivering that emergency fix if they have agile systems in place. The trick is to have a clear path for the emergency fix, clear rules for when it’s acceptable to change anything, and clear knowledge of who can authorize such changes (and people who know when to break the rules).

Article Tags

agile, Battle School, Black Friday, cloud, doorbusters, IT

About C. Thomas Tyler

C. Thomas Tyler is a Senior Consultant at Perforce Software. His career started at the NASA Kennedy Space Center in 1990. He has architected many software development environments, consulting in various organizations in financial, defense, software and lottery game development, and other sectors. He holds a B.S. in computer science from the Florida Institute of Technology.

View all posts by C. Thomas Tyler

Cookie	Duration	Description
cf_use_ob	past	Cloudflare sets this cookie to improve page load times and to disallow any security restrictions based on the visitor's IP address.
cookielawinfo-checkbox-advertisement	1 year	Set by the GDPR Cookie Consent plugin, this cookie is used to record the user consent for the cookies in the "Advertisement" category .
cookielawinfo-checkbox-analytics	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Analytics".
cookielawinfo-checkbox-functional	11 months	The cookie is set by GDPR cookie consent to record the user consent for the cookies in the category "Functional".
cookielawinfo-checkbox-necessary	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookies is used to store the user consent for the cookies in the category "Necessary".
cookielawinfo-checkbox-others	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Other.
cookielawinfo-checkbox-performance	11 months	This cookie is set by GDPR Cookie Consent plugin. The cookie is used to store the user consent for the cookies in the category "Performance".
CookieLawInfoConsent	1 year	Records the default button state of the corresponding category & the status of CCPA. It works only in coordination with the primary cookie.
JSESSIONID	session	The JSESSIONID cookie is used by New Relic to store a session identifier so that New Relic can monitor session counts for an application.
PHPSESSID	session	This cookie is native to PHP applications. The cookie is used to store and identify a users' unique session ID for the purpose of managing user session on the website. The cookie is a session cookies and is deleted when all the browser windows are closed.
viewed_cookie_policy	11 months	The cookie is set by the GDPR Cookie Consent plugin and is used to store whether or not user has consented to the use of cookies. It does not store any personal data.

Cookie	Duration	Description
__atuvc	1 year 1 month	AddThis sets this cookie to ensure that the updated count is seen when one shares a page and returns to it, before the share count cache is updated.
__atuvs	30 minutes	AddThis sets this cookie to ensure that the updated count is seen when one shares a page and returns to it, before the share count cache is updated.
__cf_bm	30 minutes	This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.

Cookie	Duration	Description
__gads	1 year 24 days	The __gads cookie, set by Google, is stored under DoubleClick domain and tracks the number of times users see an advert, measures the success of the campaign and calculates its revenue. This cookie can only be read from the domain they are set on and will not track any data while browsing through other sites.
_ga	2 years	The _ga cookie, installed by Google Analytics, calculates visitor, session and campaign data and also keeps track of site usage for the site's analytics report. The cookie stores information anonymously and assigns a randomly generated number to recognize unique visitors.
_ga_S6PB8V57DG	2 years	This cookie is installed by Google Analytics.
_gat_gtag_UA_846073_1	1 minute	Set by Google to distinguish users.
_gid	1 day	Installed by Google Analytics, _gid cookie stores information on how visitors use a website, while also creating an analytics report of the website's performance. Some of the data that are collected include the number of visitors, their source, and the pages they visit anonymously.
_jsuid	1 year	This cookie contains random number which is generated when a visitor visits the website for the first time. This cookie is used to identify the new visitors to the website.
at-rand	never	AddThis sets this cookie to track page visits, sources of traffic and share counts.
CONSENT	2 years	YouTube sets this cookie via embedded youtube-videos and registers anonymous statistical data.
iutk	5 months 27 days	This cookie is used by Issuu analytic system to gather information regarding visitor activity on Issuu products.
uvc	1 year 1 month	Set by addthis.com to determine the usage of addthis.com service.
vuid	2 years	Vimeo installs this cookie to collect tracking information by setting a unique ID to embed videos to the website.
WMF-Last-Access	1 month 14 hours 26 minutes	This cookie is used to calculate unique devices accessing the website.

Cookie	Duration	Description
__Host-GAPS	2 years	This cookie allows the website to identify a user and provide enhanced functionality and personalisation.
_pxhd	session	Used by Zoominfo to enhance customer data.
IDE	1 year 24 days	Google DoubleClick IDE cookies are used to store information about how the user uses the website to present them with relevant ads and according to the user profile.
loc	1 year 1 month	AddThis sets this geolocation cookie to help understand the location of users who share the information.
mc	1 year 1 month	Quantserve sets the mc cookie to anonymously track user behaviour on the website.
test_cookie	15 minutes	The test_cookie is set by doubleclick.net and is used to determine if the user's browser supports cookies.
VISITOR_INFO1_LIVE	5 months 27 days	A cookie set by YouTube to measure bandwidth that determines whether the user gets the new or old player interface.
YSC	session	YSC cookie is set by Youtube and is used to track the views of embedded videos on Youtube pages.
yt-remote-connected-devices	never	YouTube sets this cookie to store the video preferences of the user using embedded YouTube video.
yt-remote-device-id	never	YouTube sets this cookie to store the video preferences of the user using embedded YouTube video.
yt.innertube::nextId	never	This cookie, set by YouTube, registers a unique ID to store data on what videos from YouTube the user has seen.
yt.innertube::requests	never	This cookie, set by YouTube, registers a unique ID to store data on what videos from YouTube the user has seen.

Cookie	Duration	Description
__gpi	1 year 24 days	No description
__Secure-YEC	1 year 1 month	No description
_heatmaps_g2g_100754890	10 minutes	No description
_techvalidate_session	session	No description
cf_7166_id	20 years	No description
cf_7166_person_last_update	session	No description
f5avraaaaaaaaaaaaaaaa_session_	session	No description available.
GoogleAdServingTest	session	No description
Gyazo_cfwoker	7 years 2 months 17 days 7 hours	No description
incap_ses_451_2783402	session	No description
incap_ses_769_2783402	session	No description
loglevel	never	No description available.
m	2 years	No description available.
nlbi_2783402	session	No description
prism_252377639	1 month	No description
TS011605d9	session	No description
ustream-guest	session	No description available.
visid_incap_2783402	1 year	No description
xtc	1 year 1 month	No description

AI

AI and Software Development

Observability

Guide to Observability

CI/CD

A guide to CI/CD

Cloud Native

Cloud Native Content

Data

A Guide to Data

Test

Security Testing

Mobile

Mobile Testing

API

Sponsored by Parasoft

Performance

Load & Performance Testing

DevSecOps

A Guide to DevSecOps

Enterprise Security

A Guide to Security

Supply Chain Security

Supply Chain Security

Dev Manager

Dev Managers Content

Agile

A Guide To Agile

Value Stream

A Guide To Value Stream

Productivity

A Guide To Productivity

DevOps

DevOps Content

API

Gravitee.io

AI

AI and Software Development

Value Stream Management

A Guide To Value Stream

Guest View: Preparing for the online doorbusters

Article Tags

Subscribe to SDTimes

About C. Thomas Tyler

Related Articles

The AI productivity paradox in software engineering: Balancing efficiency and human skill retention

Plotly brings vibe coding to visual data app development

Four trends reshaping Kubernetes platform engineering

Podcast: Misconceptions around Agile in an AI world