Prod-U-Weather forecast: Seasonably cold winter results in code freezes

Prod-U-Weather forecast: Seasonably cold winter results in code freezes

Mike Wright

Checks Weather: Mid-60s°F in NYC...

It's only barely autumn. The leaves just started turning and I haven't even gone apple picking yet. Why are we talking about this? To some extent, planning for our various roadmaps is continuous. But accounting for predictable events should happen at least a couple of months out. We're just about one and a half months away from Black Friday and Cyber Monday, some of the biggest days for Button and our partners.

Chillier days ahead

Button engages in a commonly observed practice in software development: Code Freezes, specifically around major holidays or high-revenue sensitive time periods. These periods involve a seasonal peak in business activity (e.g., shopping), and many Button personnel will be out-of-office. In addition, many of our partners have similar freezes during the same periods. Thus, any production issue that arises during these periods can be both extremely costly and more challenging than usual to recover from.

Due to this elevated risk, we endeavor to avoid all non-essential modifications to our systems, including but not limited to releasing new features, enabling new partners, and fixing non-critical bugs. The goal is to minimize moving parts and general system entropy. So even when a change may seem "innocent enough" or intuitively low-risk, it is still to be avoided. We are being deliberately more conservative in our approach than usual.

To freeze or not to freeze?
Pros
Cons
Where Button lands

Button's policy for code freezes, as with most of our policies, attempts to find a healthy balance between safety and progress. Generally, code freezes are beneficial during times of elevated risk or decreased staff availability. The problems with code freezes are mostly associated with the hard deadline they represent.

Rather than impose a hard deadline by which all production deployments must be performed, we set a schedule over which a team of "Freeze Admins" become progressively stricter about the changes applied to our production environment. This group comprises Engineering management and leads. Button's production environment "slushes" over as we get closer to more sensitive times of the year, specifically Black Friday/Thanksgiving (US)/Cyber Monday and the winter holiday season.

Daily forecast posted in the holiday freeze Slack channel:

.png)

.png)

Realistically, there is still a hard deadline: the day the Freeze Admins no longer allow releases to production. The practice of the slush is to raise awareness, heighten visibility, and prevent any surprise/last-minute deployments.

Logistics

Every request to deploy to production must be approved by a Freeze Admin. In general, this group is accountable for the stability of Production and comprises Engineering Leadership as well as Engineering Team Leads.

Requests are formatted as an articulation of:

Requesting approval for these changes is relatively easy. The deploying software developer posts the above information in a dedicated Slack channel, tagging the Freeze Admins. The change is either approved, denied, or more information is requested. The approval, if given, is valid for half a day. A team of representatives from the Business and Product teams have the oversight power to request a change be deferred in the event that it would disrupt business operations.

A sample request is below (partner, application, and individual names modified/redacted):

Hello @freeze-admins - Seeking approval to push [repository]#[PR number], part two of two (link previous change) & should be the final change from me this week.

Does it work?

Yes! We have utilized a code freeze process since 2017. Overall, we have successfully prevented costly mistakes with minimal agony or disruption to roadmaps. Every year, we iterate on this process to improve what works and eliminate what doesn't.

The most important effects of doing this are:

Do you value fast-moving, high-trust engineering environments that operate at a meaningful scale in a distributed cloud environment? View our open roles and join the Button team today!