Customer work · Consumer internet

Nearbuy reduced monthly AWS cost by 30% and automated recurring issue recovery

Nearbuy wanted to move from reactive incident handling to a preventive reliability and cost program. BluePi established service levels, monitored critical AWS resources, optimized services and Spot capacity, and automated recurring defect detection and resolution.

Nearbuy reduced monthly AWS cost by 30% and automated recurring issue recovery system diagram

Nearbuy cloud cost optimization tied to service stability

Nearbuy needed to reduce recurring incidents and cloud cost without weakening the live service. BluePi used workshops and stakeholder interviews to define service requirements, then established service levels, monitoring, issue-management practices, recurring-fault automation, and regular performance and cost reviews.

01 · Shared service levels

Service levels turned stability and performance into measurable expectations for the team.

02 · Recurring-fault automation

Automation targeted repeatable detection and recovery patterns so engineers could spend less time on the same issues.

03 · Spot instance savings

AWS Spot instances and service changes contributed to a 30 percent monthly cost saving.

A review cycle connecting incidents, performance, and cost

BluePi gathered requirements through workshops and stakeholder interviews, defined service levels, put monitoring and issue management in place, automated recurring fault patterns, and reviewed performance and cost on a regular rhythm. The work improved response, stability, downtime, and monthly cost.

Agree on service measures

Service levels turned stability and performance into shared operating expectations rather than occasional escalation topics.

Remove recurring work

Automation targeted repeatable detection and recovery patterns so engineers could focus on unresolved causes.

Review cost with behavior

Resource choices were evaluated against workload and reliability requirements, protecting service while reducing waste.

Where this pattern fits

The same measurement discipline fits any live platform with repeated incidents and unclear unit cost.

Discuss this case with BluePi
Working-session promptsQuestions that define the delivery boundary

Use these prompts to decide whether the case fits your operating problem and what a first deployment should prove.

  1. Which incidents recur, and what common service, dependency, workload, or release condition connects them?
  2. What service levels matter to users, and can current monitoring show whether each one is being met?
  3. Which resources carry fixed cost for demand that is variable, interruptible, scheduled, or short-lived?
  4. Where can automated detection or recovery remove repeated operating work without hiding the underlying fault?
  5. What review cadence will connect reliability, performance, cloud cost, and planned engineering changes after the first round of improvements?

A practical first step

BluePi can begin with a joint review of recent incidents, service measures, workload patterns, and cost allocation. The output identifies immediate operating fixes, recurring work suited to automation, and architecture or capacity changes that need a measured trial.

Case details

Open a section to review the customer problem, implementation, business change, and architecture.

01The starting point

Platform support focused on urgent issues after they caused downtime or productivity loss. The team needed current health data, defined service expectations, preventive alerts, and a repeatable way to remove recurring faults.

  • System stability: Frequent issues impacted platform availability and user experience.
  • Performance: The platform needed faster response times and better efficiency.
  • Cost: Infrastructure cost needed to come down to improve return on investment (ROI).
  • Proactive approach: Operations needed to move from reactive issue resolution to early detection and prevention.
02System delivered

BluePi ran functional and non-functional requirement workshops, defined service levels and issue-management processes, monitored critical AWS resources including Spot instances, customized alerts, optimized services and capacity, and automated recurring defect detection and resolution.

  • Consultation: Workshops and stakeholder interviews identified functional and non-functional requirements.
  • Process establishment: BluePi defined service-level agreements (SLAs) and established monitoring and issue-management practices.
  • Performance improvements: Automation now detects and resolves recurring issues, reducing downtime.
  • Cost reduction: AWS Spot Instances and right-sized services cut monthly cloud cost by 30%.
03Outcomes

The program produced a 30 percent month-on-month cost saving, better performance, lower downtime, stronger stability, and less repeated engineering work.

  • 30% cost reduction: Right-sized services and Spot Instances reduced monthly cloud spending by 30%.
  • Improved performance: Response times improved and the platform became more stable.
  • Continuous improvement: Regular review meetings identify and address emerging issues.
04Operating loop

AWS service and application health entered monitoring and customized alerts. Recurring defects moved through automated detection and resolution where safe, while regular performance and cost reviews identified the next optimization decision.

05Technology estate

The platform used microservices, Docker, Kubernetes, Erlang, MongoDB, PostgreSQL, Amazon EMR, and Bamboo. The reliability program monitored the behavior of this estate instead of replacing it with a new technology stack.

06What changed for reliability and cost

The team moved from urgent incident resolution toward defined service measures, preventive alerts, automated handling of recurring faults, and a regular review rhythm for reliability and cost.

System diagram