HSPStorage Intelligence
telemetry · live
Independent platform · OceanStor Dorado V6 / V7

Storage telemetry,understood.

HSP turns performance counters, internal FTDS statistics, logs and alarms from OceanStor Dorado arrays into explained anomalies, evidence-backed diagnoses and ready-to-send reports — on premises, fully offline.

01 — Platform

One pane for everything your arrays emit.

Performance dumps, REST inventory, SSH diagnostics, SmartKit archives and live alarms flow into a single time-series and analytics stack — with every heavy task queued, retried and audited.

  1. SourcesREST API · SSH / SFTP · performance dumps · SmartKit archives · alarms
  2. CollectorsInventory sync every 5 min · alarm poller every 60 s · perf watcher · log collector
  3. StorageVictoriaMetrics time series · PostgreSQL job queue & history · object storage for artifacts
  4. IntelligencePattern detectors · ML root cause · performance envelope · health & fault engines
  5. OutputsGrafana · web console · Excel / Word reports · AI assistants via MCP
0
performance metrics decoded
across 60 resource types
0
dashboard panels, ready on day one
grafana · 154 rows
0
read-only inspection checks
database · ssh · time series
0
fault-diagnosis checks
15 categories · labs
0
performance archives, streamed
no full-unpack staging
0
alarm recommendations catalogued
plus 118 disk error codes
0
disks: health verdicts match the vendor tool
7 arrays · parity test
0
AI tools for MCP assistants
6 groups · read-only
02 — Intelligence labs

Root cause, with the reasoning shown.

A latency spike is a symptom. HSP looks at the peers around it, tests which signals moved first and how far they left their normal range, and tells you which one most likely caused it — with every number on the table.

A

Pattern detectors

FC port transmission latency, high utilisation, write flow control and baseline “find the difference” — grouped into episodes and incidents, run automatically after every upload.

B

Two-phase root cause

Peers are selected automatically, then ranked by Granger causality, permutation importance and deviation from baseline. Results in ~30 s across ~2,500 series.

C

Guard rails, not guesses

Eleven audited post-model filters, an independent “extreme movers” check and lead / lag in seconds keep the verdict honest.

D

Performance envelope

Quantile models learn each array’s IOPS, bandwidth and latency ceilings, with drift detection when the workload changes shape.

root-cause candidates · demo incidentconfidence
confidence = 0.4·granger + 0.3·baseline + 0.3·importance granger baseline importance
lead / lag · which signal moved firstΔt = 6 s
01

Detect

Envelope and pattern detectors flag a deviation the moment data lands.

02

Correlate

Findings collapse into episodes and incidents instead of an alarm storm.

03

Explain

Ranked causes with confidence, lead time and links to the exact charts.

04

Report

Engineer-grade evidence and management-ready Excel and Word reports.

03 — Proactive support

Evidence first. Before the ticket.

From scheduled health checks to offline diagnosis of a full log archive, every finding points to the line of evidence behind it — so the conversation with support starts at the answer, not at the question.

health check11 categories
0 HEALTH

SMART health score, 0–100

Per-disk scoring across 160+ drive models, fleet reports and a 16-section Word report generated from a single archive.

fault diagnosis · labs86 checks · 15 categories

Every verdict cites its source

Offline diagnosis of a collected log archive. Findings carry typed evidence down to the file and line, so engineers can verify in seconds.

findingBBU discharge test overdue on ctrl 0A
severitymajor
evidencelog · bbu_status.txt : line 418
actionschedule discharge test in maintenance window
inspection96 checks

Three speeds of “is it healthy?”

quick<100 ms
standard2–5 s
deep5–10 s

Read-only checks over the database, SSH and time series — with a 0–100 score and JSON / XML export.

deep telemetry · ftds statistics~1,200 stat files per archive

See inside the controller

HSP decodes the array’s internal FTDS statistics — hundreds of millions of points per array, published in minutes — and surfaces the noisiest IO-path locations and hottest CPU functions.

inventory & change historyevery 5 min

Know what changed, and when

Controllers, disks, ports, LUN mappings, SFPs, BBUs and PSUs synced over REST; hourly snapshots deduplicated by hash; alarms polled every 60 s across up to 32 arrays in parallel.

log collection17 file types

One click from array to report

Three collection strategies over SSH / SFTP pull exactly the archives support needs, then chain straight into health check and diagnosis.

diagnosis infotrace logsperf dumpsconfig backupsalarmscustom files
04 — AI-ready

Ask your fleet. In plain words.

A built-in Model Context Protocol server gives any MCP-capable assistant read-only access to metrics, health data and internal statistics — so an engineer can ask a question instead of building a query.

› which arrays had write latency above their envelope this week, and what led it?
0
mcp tools
05 — Built for closed networks

Runs where the internet doesn’t.

access

Identity & control

  • Local and LDAP / Active Directory sign-in
  • 3 system roles, 20 granular permissions
  • Per-array scope and full access audit
security

Hardened by default

  • Array credentials encrypted at rest
  • Strict security headers, container image scanning
  • Offline Ed25519 licensing with anti-tamper
delivery

Air-gapped by design

  • Single Docker Compose bundle, no cloud calls
  • Self-checking install, update and two-level rollback
  • Twelve-gate bundle verifier before every release
06 — Positioning

DME manages. HSP investigates.

Huawei DME is the vendor’s management and automation layer. HSP is built to go deeper on a narrower job: explaining performance and health on OceanStor Dorado. Most teams will want both.

huawei dme

Manage & automate

Huawei’s lifecycle platform for storage: provisioning with approval workflows, multi-vendor and SAN fabric management, alarm handling, AI risk prediction and forecasting — plus the DME IQ cloud service that connects arrays to Huawei support.

hsp

Explain & prove

Deep, evidence-backed analysis of what arrays are actually doing — internal statistics, logs and performance — with explainable root cause and offline operation.

CapabilityHuawei DMEHSP
Official vendor support, automatic service requestsyesvia DME IQ cloud servicenoindependent product
Provisioning & orchestrationyesservice levels, templates, approval workflowsnoread-only by design
Multi-vendor storage, SAN switches, hosts, vCenteryesnofocused on OceanStor Dorado
Capacity & performance forecastingyesrisk prediction, forward-looking forecastspartialperformance envelope (labs)
Anomaly detectionyesdynamic thresholds; log anomalies in recent releasesyespattern detectors + ML (labs)
Findings explained: thresholds, lead time, confidence breakdownnot in public docsyesevery number shown, linked to charts
Diagnosis of offline log archives & SmartKit collectionsnot in public docsSmartKit itself inspects offlineyesevidence down to file and line
Analysis of internal FTDS statisticsnot in public docsyesnoisy locations, hot functions
Telemetry stays on sitepartialDME on premises; DME IQ uploads to Huawei Cloudyes
Fully air-gapped operationnot in public docsDME IQ needs a route to Huawei Cloudyesoffline licence, no cloud calls
Open APIsyesREST, SNMP, Ansible, Kubernetes CSIyesREST API, MCP server for AI assistants

yes   partial   not in public docs — DME column based on publicly available Huawei documentation; capabilities vary by version and licence. Recent DME releases add log-anomaly detection and root-cause location. Where both tools fit, HSP complements DME rather than replacing it.