7 papers
Stochastic Connectivity as the Foundation of a Runtime Model for Microservice Availability Analysis
Anatoly A. Krasnovsky, Anna Maslovskaya
Microservice availability is commonly assessed by fault injection and chaos experiments, but such experiments are costly, operationally risky, and difficult to repeat for every arc…
Toward Operationalizing Rasmussen: Drift Observability on the Simplex for Evolving Systems
Anatoly A. Krasnovsky
Software operations increasingly rely on SLOs, traces, deployment specifications, and change events, yet dashboards and thresholding practices often expose share-like operational s…
Emergence-as-Code as a Foundation for Self-Governing Reliable Systems
Anatoly A. Krasnovsky
Service-level objective (SLO)-as-code tools make per-service reliability declarative, but users experience journeys: end-to-end executions whose availability and tail latency emerg…
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
Anatoly A. Krasnovsky
While distributed tracing and chaos engineering are becoming standard for microservices, resilience models remain largely manual and bespoke. We revisit a trace-discovered connecti…
Model Discovery and Graph Simulation: A Lightweight Gateway to Chaos Engineering
Anatoly A. Krasnovsky
Chaos engineering reveals resilience risks but is expensive and operationally risky to run broadly and often. Model-based analyses can estimate dependability, yet in practice they…
Measuring Uncertainty in Transformer Circuits with Effective Information Consistency
Anatoly A. Krasnovsky
Mechanistic interpretability has identified functional subgraphs within large language models (LLMs), known as Transformer Circuits (TCs), that appear to implement specific algorit…