NuPulse Back to nustudio.ai
Reliability and Self-Healing for Enterprise Applications

Stop getting faster
at fixing the same failure.

Recovering faster each time is progress. Never seeing it again is the point. NuPulse predicts the failure, prescribes the fix, and acts inside the limits your team sets.

Overview

Detection got good.
Nothing after it did.

Your tools can see the problem. Nothing decides what to do about it, nothing is allowed to act, and nothing removes the reason it happened. So the same failure comes back next month, and your team handles it a little faster than last time.

Speed is not the target. NuPulse prescribes the specific fix rather than describing the problem, carries it out inside the limits your team sets, and then removes the class of failure so the work does not come back.

NuPulse runs inside your SRE and operations group. Your team onboards applications, runs the agents, and sets what each one may do, using the operational access the team already has. It sits beside your existing tools and never becomes a system of record.

73%
of organizations report outages caused by alerts that were seen and ignored
Splunk, State of Observability 2025
30%
of engineering time lost to manual repetitive work
Catchpoint, SRE Report 2025
10,675
tickets a month per organization, on average
HDI, State of Tech Support 2025
2 to 10
monitoring tools per team, none reasoning across them
Catchpoint, SRE Report 2025
Control
Rules decide.
Models propose.
Your team owns both.
Trust
Nothing installed.
Nothing copied.
Nothing acts until you say so.
Structure
Relationships are declared,
not inferred.
Signal
One fault,
not forty tickets.

Maturity

Stage three tells you. Stage four does something about it.

Most enterprises have made real progress to stage three. The step after it is not a data problem.

01
Reactive

Someone reports it first.

02
Diagnostic

Anomalies detected, causes described.

03
Predictive

You know hours ahead. A person still acts.

04
Prescriptive

The specific action, costed and timed.

05
Orchestrated

Carried out across teams, inside policy.

06
Self-healing

Known classes close with nobody paged.

FOUR THINGS CROSS THAT LINE, AND EVERY ONE IS A CONTROL
01 / DECISION

Rules return the verdict. Model output cannot overrule it.

02 / ACTION PATH

Executed by the owning team's agent, under their credentials.

03 / VERIFICATION

Not done until checked, and rolled back if it did not land.

04 / RECORD

Which rule, which evidence, which system, who or what.

The loop

Predict. Detect. Diagnose. Decide.
Remediate. Verify. Learn.

Seven stages, one loop, and it gets smaller each time round. The three highlighted are where most teams have nothing at all, which is why their loop never shrinks however fast the middle runs.

01 / PREDICT

The incident that never opens

Saturation, resource drift, certificate expiry and dependency degradation forecast against the error budget, and releases scored against the incidents they historically cause.

02 / DETECT

One problem, not a queue

Alert storms collapse into a single problem with a ranked cause set. MTTD falls without new instrumentation.

03 / DIAGNOSE

Which component, which change, which owner

Evidence cited, so an engineer checks the reasoning instead of trusting it. Cross-application paths traced, not guessed.

04 / DECIDE

What lets the fix run unattended

A deterministic verdict against limits your team declared: risk tier, action limits, change window, run budget. Without it, every fix needs a person in the middle.

05 / REMEDIATE

An action, not a recommendation

The agent runs the matched runbook: scale, restart, fail over, roll back, clear the queue. This is the step that moves MTTR.

06 / VERIFY

Proof, not hope

Checked for whether it landed and whether anything downstream moved, and rolled back if not. Nothing counts as done until this passes.

07 / LEARN

Remove the class, not the instance

Occurrences clustered, the underlying class identified, a structural fix raised with evidence. Never seeing it again is the point.

WHY THE THREE ARE HIGHLIGHTED

Detect, diagnose, remediate and verify are the incident, and most teams run them well. Predict stops the loop starting, decide is what allows it to run without a person in the middle, and learn is what removes the reason it ran at all. A loop without those three can only ever get faster. With them it gets shorter, then rarer, then gone.

What you get

Seven components your team runs.

Not a black box that answers questions. A set of parts your engineers configure, inspect and switch off individually.

01 / SIGNAL INDEX

Reads and compresses before it reasons

Pulls from your platforms through their own interfaces, then indexes and compresses so an investigation costs a fraction of reading raw telemetry. This is what keeps cost flat as alert volume rises.

02 / APPLICATION MAP

Ownership and dependency, declared

Who owns what, and what depends on what. Added when a question needs more than one hop, not before. Declared rather than inferred, so it does not decay silently.

03 / RELIABILITY PLANNER

Plans the run and re-plans

Reads the signal, calls the specialists it needs, re-plans after each step, and produces costed options with cited evidence.

04 / POLICY ENGINE

The deterministic verdict

Risk tier, action limits, change window, run budget, separation of duties. Your team writes the limits. Model output cannot overrule them.

05 / AGENT REGISTRY

Every agent scoped and revocable

One place to see every agent, what it may touch, which version is live, and a kill switch on each. Least privilege by default.

06 / RESOLUTION MEMORY

What worked last time

Per application and per class, in your tenant. The reason the tenth occurrence is cheaper to handle than the first.

07 / DECISION RECORD

The trail your auditor asks for

Which rule fired, which evidence was cited, which action ran, what the result was. Retrievable per decision, not reconstructed from logs.

STATUS, PLAINLY

What is running today

The planner, registry, memory, policy chains, identity and access, model routing and audit search are in production on the NuStudio runtime. The application map and per-decision lineage retrieval are being built for NuPulse now.

Business outcomes

Every dashboard was green.
The shipment still missed the cutoff.

Uptime measures a component. Your customers experience a process. NuPulse orchestrates along the chain between them.

COMPONENT PROCESS STAGE BUSINESS OUTCOME Session service degrades One amber service among four hundred Order capture slows Twelve applications depend on that service Shipments miss the cutoff The number an operating review acts on COMPONENT VIEW An alert. Somebody decides how urgent it is, under pressure, without the far end of the chain. PROCESS VIEW Today's order intake is at risk, the owning team is named, and the fix is already costed.

Scroll the diagram sideways to read it.

PRIORITIZED BY CONSEQUENCE

Two amber services, one upstream of today's order intake. No longer a judgment call.

MEASURED IN SRE TERMS

Objectives, error budgets, burn rate, golden signals, toil, MTTD and MTTR, read from where you already define them.

REPORTED IN BUSINESS TERMS

Availability of a process, not only uptime of a component.

How it works

Read in place. Reason. Decide. Hand to the owner.

Each verified resolution becomes a known class WHAT YOU ALREADY RUN Observability and logging Tracing and metrics ITSM and ticketing On call and alerting Delivery pipelines Enterprise applications Read only. Queried in place. No second copy is taken. RELATIONSHIPS optional, added when needed Who owns what What depends on what Declared, not inferred, so it does not silently drift Resolution memory REASON, THEN DECIDE Planner plans, acts, verifies, re-plans proposes costed options holds no credentials proposes Rules engine risk tier, action limit, window verdict in milliseconds YOUR TEAM ACTS Scale Restart Fail over Roll back Run the matched runbook Inside limits your team declared Through access the team already has A kill switch on every agent Calls an application team's own automation where one exists VERIFY AND RECORD Resolved and closed Acted and recorded Handed to a person Blocked and raised Did it land? Did anything downstream move? Which rule, which evidence, which system, who or what. GOVERNANCE AND AUDIT agent identity, least privilege, decision lineage, access control, residency, read only mode, quality checks, rollback and kill switches Reference architecture, modeled and illustrative.

Scroll the diagram sideways to read it.

Architecture diagrams are illustrative. Cross system entity resolution and rule level decision lineage retrieval are specified interfaces rather than shipped capability today.

What runs where

Three techniques, and the expensive one is used least.

Most of the work across hundreds of applications is high volume classification and forecasting, which needs no language model at all. Running one there would be slower, costlier and no more accurate.

WHAT RUNS WHERE IT IS USED RULES deterministic, no model the decision to act risk tier action limits change window separation of duties run budget never probabilistic STATISTICAL AND CLASSICAL ML self hosted, retrainable anomaly detection saturation forecasting burn rate alert clustering deduplication ticket classification change risk scoring cheap and explainable at volume LANGUAGE MODELS routed per task root cause on novel incidents reading runbooks and tickets weighing options drafting the record plain language questions only where ambiguity is real

Scroll the diagram sideways to read it.

COST AT VOLUME

Cost per resolved incident does not scale linearly with alert volume, because the volume is handled by models that cost a fraction of a language model call.

THE DECISION IS NOT A MODEL

Rules return the verdict. Reproducible, auditable, and identical on the same inputs every time. Nothing probabilistic sits between diagnosis and action.

YOUR CHOICE OF MODEL

Large hosted, mid tier, or self hosted inside your own environment where data cannot leave the region. Routed per task, not chosen once.

Operating model

People and agents, side by side. No approval queue.

Your engineers and your agents share one context and one record. Autonomy is set once per application rather than approved action by action, because a queue does not scale and policy does.

01 / Assist

A person asks, it answers

Question an application, a dependency or an objective in plain language and get a grounded answer with its evidence. No query in four tools first. Then delegate the work and keep watching it.

The person is in charge.

02 / Act within policy

It decides and acts

Inside limits your team declared, it resolves the class and records what it did. Nobody is paged, because the limits were agreed once rather than per action.

This is where self-healing lives.

03 / Escalate

It hands over, with context

Novel, ambiguous or out of limits. The engineer gets a diagnosis and options, not a raw alert.

An exception path, not the main path.

Integrations

If it has an API, it connects.

Read in place, through interfaces your platforms already publish. The systems below are what customers ask about most, not a limit.

Observability and logging

Signals, metrics, traces, logs
SplunkElasticDatadog DynatraceAppDynamicsNew Relic GrafanaPrometheusOpenTelemetry LokiSumo LogicAzure Monitor CloudWatchGoogle Cloud Ops ThousandEyesSolarWindsZabbix InstanaOpenSearch

Service management and on call

Tickets, problems, changes, paging
ServiceNowJira Service Management BMC HelixFreshserviceCherwell ZendeskPagerDutyOpsgenie xMattersSlackMicrosoft Teams Statuspage

Delivery and infrastructure

Changes, deployments, platforms
GitHubGitLabAzure DevOps JenkinsArgo CDTerraform KubernetesOpenShiftVMware AWSAzureGoogle Cloud Ansible

Applications and data

Systems of record, analytical platforms
SAPSalesforceEpic GuidewireOracleDynamics SnowflakeDatabricksKafka PostgreSQLSQL Server Mainframe, via adapter
HOW IT CONNECTS

REST and GraphQL, MCP tools and servers, OpenTelemetry, Prometheus query, webhooks, syslog, JDBC and ODBC, queues, file drop. Adding a platform is configuration, not a build.

THE LEGACY HALF

Mainframe, midrange, thick client and packaged apps with no API at all, reached through a governed adapter inside your environment, under the same policy and record.

THE REAL CONSTRAINT AT SCALE

Rate limits, not model quality. NuPulse indexes and compresses before it reasons, and carries a cost budget per run, so the economics hold at hundreds of applications.

Governance

Built so an automated action can be explained.

  • Deterministic gating. Rules score risk, not the model, and cannot be overruled by it.
  • Autonomy per agent, action and environment. Set by your team per application, never uniform across all of them.
  • Separation of duties. An agent that proposes a financially controlled change cannot approve it.
  • Decision lineage. Which rule, which evidence, which approver, which system wrote what.
  • Read only first. Proposals graded against what the team actually did.
  • Rollback and kill switches. Revocable instantly, per agent, by whoever runs it.
  • Least privilege identity. Every agent scoped and revocable in one place.
WHERE IT RUNS

Your cloud, on premises, or air gapped. Signals queried in place, residency kept in region. Model choice per task, including a self hosted model where data cannot leave.

Every model call from every agent passes the same governance path, so there is no route by which a model reaches a system of record without a rule having decided first.

HOW PROGRESS IS PROVEN

Against your existing baseline: MTTD and MTTR, share of incidents closed with nobody involved, repeat rate per class, escalation rate, tickets linked per fault. False positive rate published alongside, not netted out.

Getting there

Wherever you are on the six stages, that is the starting point.

Nobody starts at stage one. Engagement begins from what already exists, measured against the baseline you already hold.

AT DETECT OR DIAGNOSE

Add the decision and the action

Your signals and models stay as they are. The rules engine, action path, verification and record are what get added.

ALREADY PREDICTING

Move to prescription

Forecasts become costed actions, and classes you already trust start resolving inside limits your teams declare.

BUILDING IT YOURSELF

Take the components, keep the product

Take the planner, rules engine, registry, memory and audit trail behind your own interface. Your platform keeps its name, roadmap and data.

HOW WE WORK

Forward deployed engineers stand it up inside your SRE group and work alongside your team. Every application starts read only, with proposals graded against what your engineers actually did.

THE TARGET STATE

Recurring incidents close without anyone being paged. Novel ones arrive already diagnosed. Engineering time goes into removing classes of failure.

FAQ

Common questions

Does NuPulse replace our observability platform?

No. It sits above and beside the tools you already run, reads them in place, and does not become the system of record.

Who runs NuPulse, and who acts?

Your SRE and operations group runs it. The team onboards applications, builds and tunes the agents, and declares what each one may do, acting through the operational access the team already holds. NuStudio does not operate it for you and does not hold credentials on your systems. Where an application team has built its own automation, NuPulse calls that instead of duplicating it.

Do application teams have to do anything?

Not to begin with. NuPulse reads what applications already emit through your existing observability, logging and ticketing platforms, so onboarding an application does not require its team to build anything. Application owners get involved when an action would touch something outside what the SRE group already operates, and at that point they set the limit once rather than approving each action.

Which systems can it read from?

Anything that exposes an API, through REST and GraphQL, MCP tools, OpenTelemetry, Prometheus query, webhooks, syslog, JDBC and ODBC, message queues or file drop. The integrations section lists what customers ask about most often and is not a limit. The only case needing different handling is a system with no API at all, reached through a governed adapter.

Is this all language models?

No, and deliberately so. Deterministic rules make every decision about whether an action may run. Statistical and classical machine learning handle the high volume work: anomaly detection, saturation and burn rate forecasting, alert clustering, deduplication, ticket classification and change risk scoring. Language models are used only where there is genuine ambiguity and unstructured evidence to reason over, such as root cause on a novel incident, reading runbooks and change notes, weighing options, and answering questions in plain language. That split is what keeps cost per resolved incident from scaling with alert volume.

What happens when an agent gets something wrong?

Read only mode first, deterministic limits on what any action may touch, verification after every action, rollback, and a full record so a bad action is explained in minutes rather than argued about for a week. Because execution is owned locally, no single component holds enough privilege to affect every application at once. We do not claim the risk is zero.

How does NuPulse relate to NuStudio?

NuPulse is a product on the NuStudio governed runtime, using the same agent execution layer, semantic layer, memory, identity and access control, model routing and decision lineage as every other NuStudio product. Choosing NuPulse does not mean adopting a separate platform.

Which industries is it built for?

Any application portfolio large enough that no single team can see across it. NuPulse was engineered to the standard regulated sectors require, including healthcare, defense, financial services and insurance, because an automated action there has to be explainable to an auditor and not merely effective. That standard is useful everywhere, and the same product runs in consumer packaged goods, retail and distribution, manufacturing, logistics, energy, telecommunications, travel and the public sector. What matters is whether business processes cross more applications than one team owns.

One product on the NuStudio runtime.

The platform pages cover the semantic layer, the agent harness, governance at the reasoning level, and how forward deployed engineers work inside your environment.