MemberPad Reliability Report: Our Uptime, Incidents, and Operational Priorities for the Year Ahead
A community platform is not a casual piece of software. When MemberPad goes down, member-to-member conversations stop. Weekly calls get rescheduled. New signups fail. Churn emails do not get processed. Payouts can stall. For a creator whose community is their livelihood, a bad 90 minutes in our infrastructure can undo a good month.
We take that responsibility seriously. We also think the best way to take it seriously is to be straight with you about what is going well, what is not, and what we are doing about it. This post is our first public reliability report. We plan to publish one every six months.
No spin, no polished euphemisms. Just the numbers and the context.
The top-line numbers
Over the last six months (November 2025 through April 2026), MemberPad's production services had:
99.982 percent overall uptime
7 minor incidents
2 major incidents
0 data-loss incidents
0 security incidents involving customer data
For comparison, 99.9 percent uptime allows roughly 8 hours 45 minutes of downtime per year. 99.99 percent allows roughly 52 minutes. Our current run puts us between those two, closer to the higher number.
We are not going to pretend we are happy with that. We are aiming for 99.99 percent or better. We will explain what stands between us and that goal below.
What we consider an incident
Before we go further, here is how we define our terms. Transparency requires agreeing on what counts.
A minor incident is a partial degradation of service, typically impacting less than 15 percent of traffic, for less than 30 minutes, where most users would not have noticed unless they were doing something specific at the time.
A major incident is a significant degradation or full outage that a reasonable user would notice and be disrupted by, affecting a large portion of customers for any meaningful duration.
We do not count planned maintenance windows as incidents, but we announce them in advance and schedule them during the lowest-traffic hours for most of our user base.
The two major incidents
Because these are the ones that matter most to you, here is the full story on each.
Major incident 1: January 17, 2026
On January 17, between 14:22 and 15:38 UTC, community feeds across all shards returned stale data. Members saw older posts, and new posts were not appearing in timelines for about 76 minutes.
Root cause: a scheduled background job for cache invalidation failed silently after a dependency update we had rolled out the night before. The cache was serving from a stale replica. Our monitoring caught the anomaly, but our initial triage misread the symptom as a database issue.
What we did: Rolled back the dependency. Rebuilt the affected cache layer. Members saw fresh content within about 10 minutes of the fix.
What we learned: Our alerting needed a distinct signal for cache freshness separate from general query performance. We have added it. We have also added an automated health check on the background job specifically. We also restructured the deployment review process so any dependency update is paired with an explicit check of the jobs that rely on it.
Apology: members during that window had a confusing 76 minutes inside their communities. We are genuinely sorry for that.
Major incident 2: March 8, 2026
On March 8, between 03:11 and 04:47 UTC, login and signup flows failed for approximately 42 percent of traffic. Existing sessions continued to work. Users trying to start a new session received errors.
Root cause: a routine database failover during a regional maintenance event took 6x longer than expected. Our automated systems correctly routed traffic to the standby, but the standby's connection pool had not been pre-warmed at the expected scale due to a configuration drift that crept in during a previous migration.
What we did: manually widened the pool. Recovered login services. Brought the primary back after the regional maintenance completed. Shipped a post-incident fix to the standby configuration to prevent drift.
What we learned: our game-day drills had not included a scenario at our current scale. They do now. We also added automatic pool-size verification to the daily health reports, which flags drift before it matters.
Apology: users who tried to log in during that 96-minute window got errors. For creators mid-signup flow, this was especially painful. We are sorry.
What you did not see
The most reliable infrastructure is usually the infrastructure whose work is invisible. Some things that ran smoothly and quietly this half:
28 hardware failures across our fleet, each routed around automatically without incident
4 certificate renewals, all automated, zero drama
61 deploys, 57 without issue, 4 that were flagged and rolled back in under 10 minutes
3 regional failovers for planned maintenance, all invisible to customers
19 security patches applied, all during automated maintenance windows
The work does not make the news. But it is what keeps the numbers above in the range they are in.
What we are prioritizing for the next six months
Here is where we are spending operational resources between now and November 2026.
Priority one: multi-region active-active posting
Right now, our post-write path is region-pinned. If a region has a bad hour, posting can be impacted for users on that region. We are building toward active-active writes so a single region can absorb the entire write load during incidents. Target: early Q4 2026.
Priority two: better observability for community-specific health
Our top-level uptime is good, but "is the signup flow healthy in this specific community right now" is a harder question to answer instantly. We are shipping better per-community health dashboards for creators to see the status of their own space. Target: Q3 2026.
Priority three: async job durability upgrades
The January 17 incident taught us that background jobs need first-class observability equal to our sync paths. We are migrating several critical jobs to a new queue system with stronger guarantees. Already underway.
Priority four: expanded game-day rehearsals
We rehearse production failures every quarter. We are expanding these to include scale-sensitive scenarios at our current peak load, not our 2024 peak. Starting this quarter.
Priority five: communications during incidents
Our status page is fine. Our in-product incident banner is better than it was a year ago. Our incident update cadence is still not where we want it. We are aiming for a clear update every 10 minutes during any major incident, no matter how redundant the update feels. Starting immediately.
What you can do to help
A few simple things that help us serve you better:
Subscribe to our status page for email or SMS alerts. You will hear from us first.
If you see a problem and your community members are messaging you, send us a quick note at [email protected]. Sometimes one motivated creator report beats our monitoring by 30 seconds.
If you ever feel like we are not being straight with you about an incident, push back. Tell us. We will either explain our reasoning or change our process. We would rather be uncomfortable with you than be unclear.
Our commitment
Every community on MemberPad is someone's serious work. Many are someone's livelihood. A few are someone's entire creative career.
We will not always be perfect. But we will always be honest. We will tell you when something went wrong, we will explain why, we will tell you what we are doing about it, and we will publish these numbers whether they make us look good or not.
Thank you for trusting us with your community. We take that trust seriously every single day.
See you back here in six months with the next report.