Hands-on Microsoft Azure monitoring and cloud support lab demonstrating Azure Monitor, Log Analytics, Azure Monitor Agent, Data Collection Rules, Kusto Query Language, performance monitoring, activity logs, alerting, Windows event investigation, controlled incident simulation, root-cause analysis, remediation, and end-to-end recovery validation.
This project simulates real-world Azure cloud support and monitoring incidents using a Windows Server virtual machine connected to Azure Monitor and Log Analytics.
The lab focuses on the operational workflow used by cloud support, systems administration, help desk, and infrastructure teams:
Collect telemetry
↓
Detect abnormal behavior
↓
Investigate with KQL
↓
Correlate guest and Azure data
↓
Identify root cause
↓
Remediate the condition
↓
Validate recovery
↓
Document the incident
The environment was intentionally designed around one monitored Windows Server so each incident could be generated, investigated, remediated, and verified from beginning to end.
Environment Setup: COMPLETE
Log Analytics Workspace: COMPLETE
Azure Monitor Agent: COMPLETE
Data Collection Rules: COMPLETE
VM Metrics Monitoring: COMPLETE
Activity Log Monitoring: COMPLETE
Azure Monitor Alerts: COMPLETE
KQL Log Analysis: COMPLETE
Incident Investigation: COMPLETE
Root-Cause Analysis: COMPLETE
Completed Incidents: 7
Documentation Files: 9
KQL Files: 11
Screenshots: 134
Name:
AZMON-WIN01
Operating System:
Windows Server 2022 Datacenter: Azure Edition
Region:
North Central US
VM Size:
Standard_D2als_v6
vCPUs:
2
Memory:
4 GiB
RG-AZMON-SUPPORT-LAB
LAW-AZMON-SUPPORT-LAB
DCR-AZMON-WINDOWS
The lab includes both:
Azure Monitor metric alerts
and:
Log Analytics scheduled query alerts
including:
ALRT-AZMON-LOW-MEMORY-KQL
with the reusable action group:
AG-AZMON-SUPPORT
flowchart TD
VM[AZMON-WIN01<br/>Windows Server 2022]
VM --> AMA[Azure Monitor Agent]
VM --> METRICS[Azure VM Metrics]
VM --> ACTIVITY[Azure Activity Log]
AMA --> DCR[DCR-AZMON-WINDOWS]
DCR --> LAW[LAW-AZMON-SUPPORT-LAB]
LAW --> PERF[Perf]
LAW --> HEARTBEAT[Heartbeat]
LAW --> EVENTS[Event Data]
LAW --> KQL[Kusto Query Language]
METRICS --> ALERTS[Azure Monitor Alerts]
KQL --> LOGALERTS[Scheduled Query Alerts]
ALERTS --> AG[AG-AZMON-SUPPORT]
LOGALERTS --> AG
PERF --> INVESTIGATION[Incident Investigation]
HEARTBEAT --> INVESTIGATION
EVENTS --> INVESTIGATION
ACTIVITY --> INVESTIGATION
INVESTIGATION --> RCA[Root-Cause Analysis]
RCA --> REMEDIATION[Remediation]
REMEDIATION --> VALIDATION[Recovery Validation]
- Microsoft Azure
- Azure Portal
- Azure Virtual Machines
- Windows Server 2022
- Azure Monitor
- Log Analytics
- Log Analytics Workspace
- Azure Monitor Agent
- Data Collection Rules
- Azure Activity Log
- Azure Metrics
- Azure Monitor Alerts
- Scheduled Query Rules
- Action Groups
- Kusto Query Language
- PowerShell
- Windows Performance Counters
- Windows Event Log
- Azure VM Run Command
- Git
- GitHub
The lab demonstrates monitoring across several Azure and operating-system layers.
Azure Resource
↓
Azure Metrics
↓
Guest Operating System
↓
Windows Performance Counters
↓
Azure Monitor Agent
↓
Data Collection Rule
↓
Log Analytics Workspace
↓
KQL Analysis
↓
Alerting
↓
Incident Response
Performance telemetry was queried from:
Perf
Examples included:
Processor utilization
Available memory
Logical disk capacity
Disk queue length
Network throughput
Azure Monitor Agent connectivity was investigated through:
Heartbeat
This allowed the lab to distinguish between:
Healthy VM + missing monitoring telemetry
and:
Actual VM availability failure
Windows Application events were generated and investigated to demonstrate:
Event creation
Event collection
KQL filtering
Event ID correlation
Severity analysis
Source analysis
Remediation validation
Azure Activity Log data was used to investigate Azure control-plane operations such as:
Virtual machine operations
Resource configuration changes
Administrative actions
Operation status
Caller information
Timestamps
The Windows monitoring environment uses:
DCR-AZMON-WINDOWS
The collected performance counters included examples such as:
Processor Information\% Processor Time
Memory\Available Bytes
LogicalDisk\Free Megabytes
LogicalDisk\Avg. Disk Queue Length
Network Interface\Bytes Total/sec
An important troubleshooting lesson from the lab was to query the actual collected performance-counter inventory before changing the DCR.
Kusto Query Language was used throughout the project to:
- Filter telemetry by computer
- Filter by performance object
- Filter by counter name
- Search recent time windows
- Find the latest sample
- Calculate minimum values
- Calculate maximum values
- Calculate averages
- Calculate deltas
- Convert bytes to MB and GB
- Classify health states
- Compare collection and ingestion timestamps
- Identify missing heartbeat
- Analyze Windows events
- Visualize incident timelines
- Build scheduled query alert conditions
- Validate post-remediation recovery
| Incident | Scenario | Primary Monitoring Area | Status |
|---|---|---|---|
| INC-001 | High CPU Utilization | Azure Metrics / Performance | Resolved |
| INC-002 | Missing Heartbeat | Azure Monitor Agent / Heartbeat | Resolved |
| INC-003 | VM Availability | Azure VM / Activity / Heartbeat | Resolved |
| INC-004 | Windows Application Event | Windows Events / Log Analytics | Resolved |
| INC-005 | Disk Space / Storage Capacity | Perf / LogicalDisk | Resolved |
| INC-006 | Memory Pressure | Perf / Memory | Resolved |
| INC-007 | KQL Scheduled Query Alert | Log Analytics / Automated Alerting | Resolved |
Documentation:
INC-001 — High CPU Utilization
The first incident established the core monitoring and incident-response workflow.
The lab generated controlled CPU utilization and investigated the resulting Azure monitoring data.
Skills demonstrated:
Azure VM monitoring
CPU metrics
Performance investigation
Alert validation
Controlled workload generation
Root-cause analysis
Remediation
Recovery verification
Incident documentation:
611 lines
Documentation:
This incident simulated interruption of Azure monitoring telemetry.
The investigation focused on determining whether:
The VM was offline
or:
The monitoring agent stopped reporting
The workflow included:
Heartbeat analysis
Agent validation
Telemetry-gap identification
Azure Monitor Agent troubleshooting
Recovery verification
Incident documentation:
1,627 lines
Documentation:
This incident simulated Azure VM unavailability through controlled deallocation.
The investigation correlated:
VM resource state
Heartbeat data
Azure Activity Log
Availability timeline
Recovery telemetry
This scenario demonstrates how Azure support engineers distinguish:
Monitoring-agent failure
from:
Actual infrastructure unavailability
Incident documentation:
2,250 lines
Documentation:
INC-004 — Windows Application Event
A controlled Windows Application event was generated on:
AZMON-WIN01
The generated event used:
Event ID:
200
Source:
AZMON-AppLab
Log:
Application
Type:
Error
The investigation demonstrated:
Windows Event generation
Event collection
Log Analytics ingestion
KQL filtering
Event-source identification
Severity validation
Timeline correlation
Incident closure
Incident documentation:
2,201 lines
Documentation:
This incident simulated storage-capacity reduction on the Windows VM.
Baseline:
C: Total:
126.45 GB
C: Free:
112.26 GB
Free:
88.78%
A controlled file was created:
C:\AZMON-Disk-Test.bin
Size:
10 GB
During the incident:
Free:
102.26 GB
Free Percentage:
80.87%
After remediation:
Free:
112.26 GB
Free Percentage:
88.78%
Log Analytics independently showed the same reduction and recovery through:
LogicalDisk\Free Megabytes
The final timechart demonstrated:
Healthy baseline
↓
Approximately 10 GB drop
↓
Controlled incident
↓
File removed
↓
Storage recovered
Incident documentation:
2,623 lines
Documentation:
This incident simulated controlled physical-memory pressure.
Baseline:
Total Memory:
3.99 GB
Available Memory:
2.79 GB
Memory Used:
30.23%
Log Analytics baseline:
Approximately 2933 MB available
A controlled PowerShell process allocated approximately:
1024 MB
Observed process working set:
Approximately 1105 MB
During pressure:
Windows Available:
1.73 GB
Memory Used:
56.59%
Log Analytics Available:
Approximately 1809–1827 MB
After remediation:
Windows Available:
2.77 GB
Memory Used:
30.78%
WorkerPresent:
False
Log Analytics:
Approximately 2915 MB
The incident also demonstrated a real troubleshooting condition.
The initial query searched for:
Memory\Available MBytes
but returned no data.
A performance-counter inventory showed the environment was actually collecting:
Memory\Available Bytes
The query was corrected without unnecessarily modifying the Data Collection Rule.
Incident documentation:
2,734 lines
Documentation:
INC-007 — Log Analytics Scheduled Query Alert
INC-007 converted KQL troubleshooting logic into automated Azure Monitor detection.
Alert rule:
ALRT-AZMON-LOW-MEMORY-KQL
Action group:
AG-AZMON-SUPPORT
Threshold:
Available memory < 2048 MB
Healthy baseline:
2911 MB
Healthy KQL result:
0 rows
Controlled incident:
1691 MB
Alert condition:
TRUE
Azure Monitor state:
FIRED
After remediation:
Windows Available:
2696 MB
Log Analytics Available:
2764 MB
KQL Result:
0 rows
Alert State:
RESOLVED
The complete lifecycle was validated:
Healthy
↓
KQL condition false
↓
Scheduled query rule enabled
↓
Controlled threshold violation
↓
KQL condition true
↓
Azure Monitor alert fired
↓
Remediation
↓
Telemetry recovery
↓
KQL condition false
↓
Alert automatically resolved
Incident documentation:
2,900 lines
The incidents were intentionally designed to increase in complexity.
INC-001
CPU Performance
↓
INC-002
Monitoring Agent / Heartbeat
↓
INC-003
Infrastructure Availability
↓
INC-004
Windows Event Investigation
↓
INC-005
Storage Capacity
↓
INC-006
Memory Capacity
↓
INC-007
Automated KQL Alerting
The project contains dedicated documentation for the major monitoring components.
Documentation/01-Environment-Setup.md
Covers:
Azure environment
Resource group
Virtual machine
Monitoring design
Lab architecture
Documentation/02-Log-Analytics-Workspace.md
Covers:
Workspace creation
Workspace configuration
Log storage
Query environment
Monitoring integration
Documentation/03-Azure-Monitor-Agent-DCR.md
Covers:
Azure Monitor Agent
Agent deployment
Data Collection Rules
Performance counters
Telemetry collection
Documentation/04-VM-Metrics-Monitoring.md
Covers:
CPU metrics
VM monitoring
Metric charts
Performance interpretation
Documentation/05-Activity-Logs-Diagnostics.md
Covers:
Azure Activity Log
Administrative operations
Resource changes
Operation status
Diagnostic investigation
Documentation/06-Azure-Monitor-Alerts.md
Covers:
Alert rules
Conditions
Thresholds
Action groups
Severity
Alert lifecycle
Documentation/07-KQL-Log-Analysis.md
Covers:
Kusto Query Language
Filtering
Aggregation
Time windows
Telemetry analysis
Operational queries
Documentation/08-Incident-Investigation.md
Covers:
Support workflow
Evidence collection
Telemetry correlation
Troubleshooting
Incident validation
Documentation/09-Root-Cause-Analysis.md
Covers:
Symptom identification
Evidence correlation
Root-cause isolation
Remediation
Recovery validation
Incident closure
Reusable KQL queries are stored in:
KQL/
The project currently contains:
11 KQL files
Used for Azure control-plane operation investigation.
Used for:
Logical disk free space
Capacity investigation
Storage incident timelines
Recovery validation
Used for general Azure Monitor Agent heartbeat analysis.
Incident-Investigation-Queries.kql
Reusable investigation queries for cross-incident troubleshooting.
Used for:
Low-memory detection
State classification
Scheduled query alerts
Recovery verification
Used for:
Available memory
Memory timelines
Minimum and maximum memory
Pressure detection
Recovery validation
Used specifically for the monitoring interruption incident.
Reusable VM performance queries.
Used to investigate:
VM availability
Heartbeat state
Controlled deallocation
Recovery
Windows-Event-Incident-Queries.kql
Used for the controlled Windows event incident.
Reusable Windows event investigation queries.
Perf
| where TimeGenerated > ago(30m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| top 1 by TimeGenerated desc
| project TimeGenerated,
Computer,
AvailableMB=round(CounterValue / 1024.0 / 1024.0,0),
AvailableGB=round(CounterValue / 1024.0 / 1024.0 / 1024.0,2)Perf
| where TimeGenerated > ago(5m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| summarize arg_max(TimeGenerated, CounterValue) by Computer
| extend AvailableMB=CounterValue / 1024.0 / 1024.0
| where AvailableMB < 2048
| project TimeGenerated,
Computer,
AvailableMB=round(AvailableMB,0)This query was used as the basis for:
ALRT-AZMON-LOW-MEMORY-KQL
Perf
| where TimeGenerated > ago(30m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| extend AvailableMB=CounterValue / 1024.0 / 1024.0
| extend State=iff(AvailableMB < 2048, "LOW MEMORY", "HEALTHY")
| project TimeGenerated,
Computer,
AvailableMB=round(AvailableMB,0),
State
| order by TimeGenerated descPerf
| where TimeGenerated > ago(60m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| summarize AvailableMB=avg(CounterValue) / 1024.0 / 1024.0
by bin(TimeGenerated,1m)
| order by TimeGenerated asc
| render timechartThe repository contains:
134 screenshots
Evidence includes:
Azure resource configuration
Log Analytics workspace configuration
Azure Monitor Agent onboarding
Data Collection Rule configuration
VM metrics
Activity logs
Alert rules
Action groups
KQL results
PowerShell validation
Incident generation
Incident timelines
Fired alerts
Remediation
Recovery validation
Resolved alerts
Screenshots are organized by monitoring phase and incident under:
Screenshots/
A typical incident contains evidence for:
01 — Healthy baseline
↓
02 — Detection logic
↓
03 — Incident generation
↓
04 — Local validation
↓
05 — Azure telemetry confirmation
↓
06 — Timeline visualization
↓
07 — Root-cause identification
↓
08 — Remediation
↓
09 — Recovery validation
↓
10 — Incident closure
The lab follows a repeatable troubleshooting method.
Determine the expected healthy state.
Examples:
Normal CPU
Normal memory
Normal disk capacity
Recent heartbeat
Healthy VM availability
Expected event volume
Validate that the reported condition actually exists.
Determine whether the issue exists at:
Azure resource layer
Guest OS layer
Agent layer
Data Collection Rule layer
Log Analytics layer
Alerting layer
Use KQL and Azure monitoring data to isolate the abnormal behavior.
Where possible, compare:
Windows
Azure Metrics
Heartbeat
Perf
Activity Log
Azure Alerts
Separate:
symptom
from:
underlying cause
Perform the lowest-risk action appropriate to the identified cause.
Do not close the incident immediately after the remediation command succeeds.
Confirm:
Guest recovery
Telemetry recovery
Query recovery
Alert recovery
Record:
What happened
What evidence confirmed it
What caused it
What was changed
How recovery was verified
During memory investigation, the initial KQL query returned no records.
The actual issue was:
Incorrect counter name
not:
Broken Azure Monitor Agent
The existing Perf data was queried before modifying the Data Collection Rule.
This avoided an unnecessary configuration change.
A missing heartbeat may indicate:
Agent interruption
while an unavailable VM may indicate:
Actual infrastructure state change
These conditions require different troubleshooting paths.
Immediately after remediation, Azure may still display an older state.
Support engineers should understand the path:
Resource changes
↓
Agent collects sample
↓
Telemetry transmitted
↓
Workspace ingests sample
↓
KQL sees new state
↓
Alert engine reevaluates
INC-007 demonstrated that the Windows server recovered before Azure Monitor displayed:
Resolved
This is expected behavior with scheduled query alert evaluation.
Multiple instances of the same alert rule can exist.
The current incident should be correlated using:
Alert timestamp
Detection timestamp
Remediation timestamp
Recovery timestamp
Resolved timestamp
The lab intentionally generated controlled faults to validate monitoring.
Examples included:
CPU pressure
Monitoring interruption
VM deallocation
Windows application error
Disk-space reduction
Memory pressure
Low-memory alert condition
Controlled tests were designed to be:
Temporary
Reversible
Observable
Safe for the lab environment
Easy to validate
flowchart LR
A[Healthy Baseline] --> B[Incident Generated]
B --> C[Telemetry Changes]
C --> D[KQL Investigation]
D --> E[Root Cause Identified]
E --> F[Remediation]
F --> G[Telemetry Recovery]
G --> H[Alert / Query Recovery]
H --> I[Incident Closed]
- Azure resource groups
- Azure virtual machines
- Azure Portal
- Azure resource monitoring
- Azure Activity Log
- Azure Monitor
- VM metrics
- Log Analytics
- Azure Monitor Agent
- Data Collection Rules
- Windows performance counters
- Heartbeat monitoring
- Metric alerts
- Scheduled query alerts
- Alert rules
- Action groups
- Severity levels
- Threshold configuration
- Automatic resolution
- Alert lifecycle investigation
whereprojectextendsummarizearg_maxminmaxavgcountiffagobinorder bytoprender timechartingestion_time()
- Windows Server 2022
- PowerShell
- CIM
- Windows Event Log
- Performance counters
- Process investigation
- CPU utilization
- Memory utilization
- Storage capacity
- VM Run Command
- Incident triage
- Baseline analysis
- Performance troubleshooting
- Agent troubleshooting
- Availability troubleshooting
- Event investigation
- Capacity troubleshooting
- Root-cause analysis
- Evidence correlation
- Recovery validation
- Incident documentation
- Evidence management
- Alert correlation
- Monitoring validation
- Remediation planning
- Incident closure
- Git version control
- GitHub documentation
Azure-Monitor-Log-Analytics-Support-Lab/
│
├── Documentation/
│ ├── 01-Environment-Setup.md
│ ├── 02-Log-Analytics-Workspace.md
│ ├── 03-Azure-Monitor-Agent-DCR.md
│ ├── 04-VM-Metrics-Monitoring.md
│ ├── 05-Activity-Logs-Diagnostics.md
│ ├── 06-Azure-Monitor-Alerts.md
│ ├── 07-KQL-Log-Analysis.md
│ ├── 08-Incident-Investigation.md
│ └── 09-Root-Cause-Analysis.md
│
├── Help-Desk-Tickets/
│ ├── INC-001-High-CPU.md
│ ├── INC-002-Missing-Heartbeat.md
│ ├── INC-003-VM-Availability.md
│ ├── INC-004-Windows-Event.md
│ ├── INC-005-Disk-Space.md
│ ├── INC-006-Memory-Pressure.md
│ └── INC-007-Log-Query-Alert.md
│
├── KQL/
│ ├── Activity-Log-Queries.kql
│ ├── Disk-Space-Queries.kql
│ ├── Heartbeat-Queries.kql
│ ├── Incident-Investigation-Queries.kql
│ ├── Log-Query-Alert-Queries.kql
│ ├── Memory-Pressure-Queries.kql
│ ├── Missing-Heartbeat-Queries.kql
│ ├── Performance-Queries.kql
│ ├── VM-Availability-Queries.kql
│ ├── Windows-Event-Incident-Queries.kql
│ └── Windows-Event-Queries.kql
│
├── Screenshots/
│ └── 134 monitoring and incident screenshots
│
└── README.md
This project was built to demonstrate practical skills relevant to roles such as:
Azure Support Engineer
Cloud Support Specialist
Technical Support Engineer
IT Support Specialist
Systems Administrator
Cloud Administrator
Help Desk Technician
Infrastructure Support Technician
Microsoft Support Specialist
NOC / Monitoring Analyst
Rather than documenting only successful Azure resource creation, the project emphasizes:
something breaks
↓
support investigates
↓
evidence is gathered
↓
root cause is identified
↓
the issue is fixed
↓
recovery is proven
Azure Monitor provides monitoring and observability across Azure resources, applications, operating systems, metrics, logs, and alerts.
In this lab I used Azure Monitor to investigate:
CPU
VM availability
memory
storage
activity logs
agent health
alerts
Log Analytics provides a query environment for analyzing monitoring data stored in an Azure Log Analytics workspace.
I used KQL to investigate:
Heartbeat
Perf
Windows events
resource state
memory
disk capacity
incident timelines
Azure Monitor Agent collects guest operating-system monitoring data and forwards it according to Data Collection Rules.
In this environment it collected Windows telemetry from:
AZMON-WIN01
into:
LAW-AZMON-SUPPORT-LAB
A Data Collection Rule determines what monitoring data should be collected and where it should be sent.
The lab used:
DCR-AZMON-WINDOWS
to collect Windows performance data.
I first determined whether the VM itself was unavailable or whether only the monitoring telemetry had stopped.
I used:
Heartbeat
VM state
Azure Monitor Agent status
Activity Log
recent telemetry timestamps
to isolate the problem.
I compared:
Windows physical memory
process working set
Memory\Available Bytes
Log Analytics Perf data
KQL timelines
The controlled worker consumed approximately 1 GB and the reduction was visible both locally and in Log Analytics.
I built a scheduled KQL query rule that detected when:
Available memory < 2048 MB
I validated:
healthy state
alert condition false
controlled threshold violation
alert firing
remediation
telemetry recovery
condition clearing
automatic alert resolution
The completed lab demonstrates the ability to:
Deploy monitoring
Collect telemetry
Query operational data
Build alerts
Simulate incidents
Investigate symptoms
Identify root causes
Remediate failures
Validate recovery
Document incidents
The final environment contains:
9 monitoring documentation files
7 completed support incidents
11 reusable KQL files
134 screenshots
Azure metric monitoring
Log Analytics monitoring
Azure Monitor Agent telemetry
Data Collection Rules
Azure Activity Log investigation
Metric alerting
KQL scheduled query alerting
Windows Server troubleshooting
End-to-end incident validation
Azure Monitor Environment:
COMPLETE
Log Analytics Environment:
COMPLETE
Azure Monitor Agent:
VERIFIED
Data Collection Rule:
VERIFIED
Performance Monitoring:
VERIFIED
Heartbeat Monitoring:
VERIFIED
Activity Log Monitoring:
VERIFIED
Windows Event Monitoring:
VERIFIED
Metric Alerting:
VERIFIED
KQL Scheduled Query Alerting:
VERIFIED
Incident Investigations:
7 COMPLETE
Reusable KQL Library:
11 FILES
Screenshot Evidence:
134
Repository Status:
PORTFOLIO READY


