Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Azure Monitor + Log Analytics Support Lab

Azure Monitor + Log Analytics Support Lab

Hands-on Microsoft Azure monitoring and cloud support lab demonstrating Azure Monitor, Log Analytics, Azure Monitor Agent, Data Collection Rules, Kusto Query Language, performance monitoring, activity logs, alerting, Windows event investigation, controlled incident simulation, root-cause analysis, remediation, and end-to-end recovery validation.


Project Overview

This project simulates real-world Azure cloud support and monitoring incidents using a Windows Server virtual machine connected to Azure Monitor and Log Analytics.

The lab focuses on the operational workflow used by cloud support, systems administration, help desk, and infrastructure teams:

Collect telemetry
      ↓
Detect abnormal behavior
      ↓
Investigate with KQL
      ↓
Correlate guest and Azure data
      ↓
Identify root cause
      ↓
Remediate the condition
      ↓
Validate recovery
      ↓
Document the incident

The environment was intentionally designed around one monitored Windows Server so each incident could be generated, investigated, remediated, and verified from beginning to end.


Lab Status

Environment Setup:        COMPLETE
Log Analytics Workspace:  COMPLETE
Azure Monitor Agent:      COMPLETE
Data Collection Rules:    COMPLETE
VM Metrics Monitoring:    COMPLETE
Activity Log Monitoring:  COMPLETE
Azure Monitor Alerts:     COMPLETE
KQL Log Analysis:         COMPLETE
Incident Investigation:   COMPLETE
Root-Cause Analysis:      COMPLETE

Completed Incidents:      7
Documentation Files:      9
KQL Files:                11
Screenshots:              134

Environment

Azure Virtual Machine

Name:
AZMON-WIN01

Operating System:
Windows Server 2022 Datacenter: Azure Edition

Region:
North Central US

VM Size:
Standard_D2als_v6

vCPUs:
2

Memory:
4 GiB

Resource Group

RG-AZMON-SUPPORT-LAB

Log Analytics Workspace

LAW-AZMON-SUPPORT-LAB

Data Collection Rule

DCR-AZMON-WINDOWS

Alerting

The lab includes both:

Azure Monitor metric alerts

and:

Log Analytics scheduled query alerts

including:

ALRT-AZMON-LOW-MEMORY-KQL

with the reusable action group:

AG-AZMON-SUPPORT

Monitoring Architecture

Azure Monitor + Log Analytics Architecture

flowchart TD
    VM[AZMON-WIN01<br/>Windows Server 2022]

    VM --> AMA[Azure Monitor Agent]
    VM --> METRICS[Azure VM Metrics]
    VM --> ACTIVITY[Azure Activity Log]

    AMA --> DCR[DCR-AZMON-WINDOWS]
    DCR --> LAW[LAW-AZMON-SUPPORT-LAB]

    LAW --> PERF[Perf]
    LAW --> HEARTBEAT[Heartbeat]
    LAW --> EVENTS[Event Data]
    LAW --> KQL[Kusto Query Language]

    METRICS --> ALERTS[Azure Monitor Alerts]
    KQL --> LOGALERTS[Scheduled Query Alerts]

    ALERTS --> AG[AG-AZMON-SUPPORT]
    LOGALERTS --> AG

    PERF --> INVESTIGATION[Incident Investigation]
    HEARTBEAT --> INVESTIGATION
    EVENTS --> INVESTIGATION
    ACTIVITY --> INVESTIGATION

    INVESTIGATION --> RCA[Root-Cause Analysis]
    RCA --> REMEDIATION[Remediation]
    REMEDIATION --> VALIDATION[Recovery Validation]
Loading

Technologies Used

  • Microsoft Azure
  • Azure Portal
  • Azure Virtual Machines
  • Windows Server 2022
  • Azure Monitor
  • Log Analytics
  • Log Analytics Workspace
  • Azure Monitor Agent
  • Data Collection Rules
  • Azure Activity Log
  • Azure Metrics
  • Azure Monitor Alerts
  • Scheduled Query Rules
  • Action Groups
  • Kusto Query Language
  • PowerShell
  • Windows Performance Counters
  • Windows Event Log
  • Azure VM Run Command
  • Git
  • GitHub

Core Monitoring Workflow

The lab demonstrates monitoring across several Azure and operating-system layers.

Azure Resource
      ↓
Azure Metrics
      ↓
Guest Operating System
      ↓
Windows Performance Counters
      ↓
Azure Monitor Agent
      ↓
Data Collection Rule
      ↓
Log Analytics Workspace
      ↓
KQL Analysis
      ↓
Alerting
      ↓
Incident Response

Data Sources Investigated

Performance Data

Performance telemetry was queried from:

Perf

Examples included:

Processor utilization
Available memory
Logical disk capacity
Disk queue length
Network throughput

Agent Health

Azure Monitor Agent connectivity was investigated through:

Heartbeat

This allowed the lab to distinguish between:

Healthy VM + missing monitoring telemetry

and:

Actual VM availability failure

Windows Events

Windows Application events were generated and investigated to demonstrate:

Event creation
Event collection
KQL filtering
Event ID correlation
Severity analysis
Source analysis
Remediation validation

Azure Activity Logs

Azure Activity Log data was used to investigate Azure control-plane operations such as:

Virtual machine operations
Resource configuration changes
Administrative actions
Operation status
Caller information
Timestamps

Data Collection Rule

The Windows monitoring environment uses:

DCR-AZMON-WINDOWS

The collected performance counters included examples such as:

Processor Information\% Processor Time
Memory\Available Bytes
LogicalDisk\Free Megabytes
LogicalDisk\Avg. Disk Queue Length
Network Interface\Bytes Total/sec

An important troubleshooting lesson from the lab was to query the actual collected performance-counter inventory before changing the DCR.


KQL Workflow

Kusto Query Language was used throughout the project to:

  • Filter telemetry by computer
  • Filter by performance object
  • Filter by counter name
  • Search recent time windows
  • Find the latest sample
  • Calculate minimum values
  • Calculate maximum values
  • Calculate averages
  • Calculate deltas
  • Convert bytes to MB and GB
  • Classify health states
  • Compare collection and ingestion timestamps
  • Identify missing heartbeat
  • Analyze Windows events
  • Visualize incident timelines
  • Build scheduled query alert conditions
  • Validate post-remediation recovery

Completed Incident Portfolio

Incident Scenario Primary Monitoring Area Status
INC-001 High CPU Utilization Azure Metrics / Performance Resolved
INC-002 Missing Heartbeat Azure Monitor Agent / Heartbeat Resolved
INC-003 VM Availability Azure VM / Activity / Heartbeat Resolved
INC-004 Windows Application Event Windows Events / Log Analytics Resolved
INC-005 Disk Space / Storage Capacity Perf / LogicalDisk Resolved
INC-006 Memory Pressure Perf / Memory Resolved
INC-007 KQL Scheduled Query Alert Log Analytics / Automated Alerting Resolved

INC-001 — High CPU Utilization

Documentation:

INC-001 — High CPU Utilization

The first incident established the core monitoring and incident-response workflow.

The lab generated controlled CPU utilization and investigated the resulting Azure monitoring data.

Skills demonstrated:

Azure VM monitoring
CPU metrics
Performance investigation
Alert validation
Controlled workload generation
Root-cause analysis
Remediation
Recovery verification

Incident documentation:

611 lines

INC-002 — Missing Heartbeat

Documentation:

INC-002 — Missing Heartbeat

This incident simulated interruption of Azure monitoring telemetry.

The investigation focused on determining whether:

The VM was offline

or:

The monitoring agent stopped reporting

The workflow included:

Heartbeat analysis
Agent validation
Telemetry-gap identification
Azure Monitor Agent troubleshooting
Recovery verification

Incident documentation:

1,627 lines

INC-003 — Virtual Machine Availability

Documentation:

INC-003 — VM Availability

This incident simulated Azure VM unavailability through controlled deallocation.

The investigation correlated:

VM resource state
Heartbeat data
Azure Activity Log
Availability timeline
Recovery telemetry

This scenario demonstrates how Azure support engineers distinguish:

Monitoring-agent failure

from:

Actual infrastructure unavailability

Incident documentation:

2,250 lines

INC-004 — Windows Application Event Investigation

Documentation:

INC-004 — Windows Application Event

A controlled Windows Application event was generated on:

AZMON-WIN01

The generated event used:

Event ID:
200

Source:
AZMON-AppLab

Log:
Application

Type:
Error

The investigation demonstrated:

Windows Event generation
Event collection
Log Analytics ingestion
KQL filtering
Event-source identification
Severity validation
Timeline correlation
Incident closure

Incident documentation:

2,201 lines

INC-005 — Disk Space / Storage Capacity Monitoring

Documentation:

INC-005 — Disk Space

This incident simulated storage-capacity reduction on the Windows VM.

Baseline:

C: Total:
126.45 GB

C: Free:
112.26 GB

Free:
88.78%

A controlled file was created:

C:\AZMON-Disk-Test.bin

Size:

10 GB

During the incident:

Free:
102.26 GB

Free Percentage:
80.87%

After remediation:

Free:
112.26 GB

Free Percentage:
88.78%

Log Analytics independently showed the same reduction and recovery through:

LogicalDisk\Free Megabytes

The final timechart demonstrated:

Healthy baseline
      ↓
Approximately 10 GB drop
      ↓
Controlled incident
      ↓
File removed
      ↓
Storage recovered

Incident documentation:

2,623 lines

INC-006 — Memory Pressure / High Memory Usage

Documentation:

INC-006 — Memory Pressure

This incident simulated controlled physical-memory pressure.

Baseline:

Total Memory:
3.99 GB

Available Memory:
2.79 GB

Memory Used:
30.23%

Log Analytics baseline:

Approximately 2933 MB available

A controlled PowerShell process allocated approximately:

1024 MB

Observed process working set:

Approximately 1105 MB

During pressure:

Windows Available:
1.73 GB

Memory Used:
56.59%

Log Analytics Available:
Approximately 1809–1827 MB

After remediation:

Windows Available:
2.77 GB

Memory Used:
30.78%

WorkerPresent:
False

Log Analytics:
Approximately 2915 MB

The incident also demonstrated a real troubleshooting condition.

The initial query searched for:

Memory\Available MBytes

but returned no data.

A performance-counter inventory showed the environment was actually collecting:

Memory\Available Bytes

The query was corrected without unnecessarily modifying the Data Collection Rule.

Incident documentation:

2,734 lines

INC-007 — Log Analytics Scheduled Query Alert / Low Memory

Documentation:

INC-007 — Log Analytics Scheduled Query Alert

INC-007 converted KQL troubleshooting logic into automated Azure Monitor detection.

Alert rule:

ALRT-AZMON-LOW-MEMORY-KQL

Action group:

AG-AZMON-SUPPORT

Threshold:

Available memory < 2048 MB

Healthy baseline:

2911 MB

Healthy KQL result:

0 rows

Controlled incident:

1691 MB

Alert condition:

TRUE

Azure Monitor state:

FIRED

After remediation:

Windows Available:
2696 MB

Log Analytics Available:
2764 MB

KQL Result:
0 rows

Alert State:
RESOLVED

The complete lifecycle was validated:

Healthy
      ↓
KQL condition false
      ↓
Scheduled query rule enabled
      ↓
Controlled threshold violation
      ↓
KQL condition true
      ↓
Azure Monitor alert fired
      ↓
Remediation
      ↓
Telemetry recovery
      ↓
KQL condition false
      ↓
Alert automatically resolved

Incident documentation:

2,900 lines

Incident Progression

The incidents were intentionally designed to increase in complexity.

INC-001
CPU Performance
      ↓
INC-002
Monitoring Agent / Heartbeat
      ↓
INC-003
Infrastructure Availability
      ↓
INC-004
Windows Event Investigation
      ↓
INC-005
Storage Capacity
      ↓
INC-006
Memory Capacity
      ↓
INC-007
Automated KQL Alerting

Documentation

The project contains dedicated documentation for the major monitoring components.

01 — Environment Setup

Documentation/01-Environment-Setup.md

Covers:

Azure environment
Resource group
Virtual machine
Monitoring design
Lab architecture

02 — Log Analytics Workspace

Documentation/02-Log-Analytics-Workspace.md

Covers:

Workspace creation
Workspace configuration
Log storage
Query environment
Monitoring integration

03 — Azure Monitor Agent + Data Collection Rule

Documentation/03-Azure-Monitor-Agent-DCR.md

Covers:

Azure Monitor Agent
Agent deployment
Data Collection Rules
Performance counters
Telemetry collection

04 — VM Metrics Monitoring

Documentation/04-VM-Metrics-Monitoring.md

Covers:

CPU metrics
VM monitoring
Metric charts
Performance interpretation

05 — Activity Logs + Diagnostics

Documentation/05-Activity-Logs-Diagnostics.md

Covers:

Azure Activity Log
Administrative operations
Resource changes
Operation status
Diagnostic investigation

06 — Azure Monitor Alerts

Documentation/06-Azure-Monitor-Alerts.md

Covers:

Alert rules
Conditions
Thresholds
Action groups
Severity
Alert lifecycle

07 — KQL Log Analysis

Documentation/07-KQL-Log-Analysis.md

Covers:

Kusto Query Language
Filtering
Aggregation
Time windows
Telemetry analysis
Operational queries

08 — Incident Investigation

Documentation/08-Incident-Investigation.md

Covers:

Support workflow
Evidence collection
Telemetry correlation
Troubleshooting
Incident validation

09 — Root-Cause Analysis

Documentation/09-Root-Cause-Analysis.md

Covers:

Symptom identification
Evidence correlation
Root-cause isolation
Remediation
Recovery validation
Incident closure

KQL Query Library

Reusable KQL queries are stored in:

KQL/

The project currently contains:

11 KQL files

Activity Log Queries

Activity-Log-Queries.kql

Used for Azure control-plane operation investigation.


Disk Space Queries

Disk-Space-Queries.kql

Used for:

Logical disk free space
Capacity investigation
Storage incident timelines
Recovery validation

Heartbeat Queries

Heartbeat-Queries.kql

Used for general Azure Monitor Agent heartbeat analysis.


Incident Investigation Queries

Incident-Investigation-Queries.kql

Reusable investigation queries for cross-incident troubleshooting.


Log Query Alert Queries

Log-Query-Alert-Queries.kql

Used for:

Low-memory detection
State classification
Scheduled query alerts
Recovery verification

Memory Pressure Queries

Memory-Pressure-Queries.kql

Used for:

Available memory
Memory timelines
Minimum and maximum memory
Pressure detection
Recovery validation

Missing Heartbeat Queries

Missing-Heartbeat-Queries.kql

Used specifically for the monitoring interruption incident.


Performance Queries

Performance-Queries.kql

Reusable VM performance queries.


VM Availability Queries

VM-Availability-Queries.kql

Used to investigate:

VM availability
Heartbeat state
Controlled deallocation
Recovery

Windows Event Incident Queries

Windows-Event-Incident-Queries.kql

Used for the controlled Windows event incident.


Windows Event Queries

Windows-Event-Queries.kql

Reusable Windows event investigation queries.


Example KQL — Latest Available Memory

Perf
| where TimeGenerated > ago(30m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| top 1 by TimeGenerated desc
| project TimeGenerated,
          Computer,
          AvailableMB=round(CounterValue / 1024.0 / 1024.0,0),
          AvailableGB=round(CounterValue / 1024.0 / 1024.0 / 1024.0,2)

Example KQL — Low-Memory Detection

Perf
| where TimeGenerated > ago(5m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| summarize arg_max(TimeGenerated, CounterValue) by Computer
| extend AvailableMB=CounterValue / 1024.0 / 1024.0
| where AvailableMB < 2048
| project TimeGenerated,
          Computer,
          AvailableMB=round(AvailableMB,0)

This query was used as the basis for:

ALRT-AZMON-LOW-MEMORY-KQL

Example KQL — Memory Health Classification

Perf
| where TimeGenerated > ago(30m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| extend AvailableMB=CounterValue / 1024.0 / 1024.0
| extend State=iff(AvailableMB < 2048, "LOW MEMORY", "HEALTHY")
| project TimeGenerated,
          Computer,
          AvailableMB=round(AvailableMB,0),
          State
| order by TimeGenerated desc

Example KQL — Incident Timeline

Perf
| where TimeGenerated > ago(60m)
| where Computer =~ "AZMON-WIN01"
| where ObjectName =~ "Memory"
| where CounterName =~ "Available Bytes"
| summarize AvailableMB=avg(CounterValue) / 1024.0 / 1024.0
    by bin(TimeGenerated,1m)
| order by TimeGenerated asc
| render timechart

Screenshot Evidence

The repository contains:

134 screenshots

Evidence includes:

Azure resource configuration
Log Analytics workspace configuration
Azure Monitor Agent onboarding
Data Collection Rule configuration
VM metrics
Activity logs
Alert rules
Action groups
KQL results
PowerShell validation
Incident generation
Incident timelines
Fired alerts
Remediation
Recovery validation
Resolved alerts

Screenshots are organized by monitoring phase and incident under:

Screenshots/

Example Incident Evidence Flow

A typical incident contains evidence for:

01 — Healthy baseline
      ↓
02 — Detection logic
      ↓
03 — Incident generation
      ↓
04 — Local validation
      ↓
05 — Azure telemetry confirmation
      ↓
06 — Timeline visualization
      ↓
07 — Root-cause identification
      ↓
08 — Remediation
      ↓
09 — Recovery validation
      ↓
10 — Incident closure

Support Troubleshooting Methodology

The lab follows a repeatable troubleshooting method.

1. Establish Baseline

Determine the expected healthy state.

Examples:

Normal CPU
Normal memory
Normal disk capacity
Recent heartbeat
Healthy VM availability
Expected event volume

2. Confirm the Symptom

Validate that the reported condition actually exists.


3. Identify the Monitoring Layer

Determine whether the issue exists at:

Azure resource layer
Guest OS layer
Agent layer
Data Collection Rule layer
Log Analytics layer
Alerting layer

4. Query the Evidence

Use KQL and Azure monitoring data to isolate the abnormal behavior.


5. Correlate Multiple Sources

Where possible, compare:

Windows
Azure Metrics
Heartbeat
Perf
Activity Log
Azure Alerts

6. Identify Root Cause

Separate:

symptom

from:

underlying cause

7. Remediate

Perform the lowest-risk action appropriate to the identified cause.


8. Validate Recovery

Do not close the incident immediately after the remediation command succeeds.

Confirm:

Guest recovery
Telemetry recovery
Query recovery
Alert recovery

9. Document Closure

Record:

What happened
What evidence confirmed it
What caused it
What was changed
How recovery was verified

Key Troubleshooting Lessons

Query No Results Does Not Always Mean Monitoring Is Broken

During memory investigation, the initial KQL query returned no records.

The actual issue was:

Incorrect counter name

not:

Broken Azure Monitor Agent

Inventory Telemetry Before Changing Configuration

The existing Perf data was queried before modifying the Data Collection Rule.

This avoided an unnecessary configuration change.


Monitoring and Resource State Are Different

A missing heartbeat may indicate:

Agent interruption

while an unavailable VM may indicate:

Actual infrastructure state change

These conditions require different troubleshooting paths.


Telemetry Has Collection and Ingestion Delay

Immediately after remediation, Azure may still display an older state.

Support engineers should understand the path:

Resource changes
      ↓
Agent collects sample
      ↓
Telemetry transmitted
      ↓
Workspace ingests sample
      ↓
KQL sees new state
      ↓
Alert engine reevaluates

Alert Resolution Can Lag Resource Recovery

INC-007 demonstrated that the Windows server recovered before Azure Monitor displayed:

Resolved

This is expected behavior with scheduled query alert evaluation.


Correlate the Correct Alert Instance

Multiple instances of the same alert rule can exist.

The current incident should be correlated using:

Alert timestamp
Detection timestamp
Remediation timestamp
Recovery timestamp
Resolved timestamp

Controlled Incident Philosophy

The lab intentionally generated controlled faults to validate monitoring.

Examples included:

CPU pressure
Monitoring interruption
VM deallocation
Windows application error
Disk-space reduction
Memory pressure
Low-memory alert condition

Controlled tests were designed to be:

Temporary
Reversible
Observable
Safe for the lab environment
Easy to validate

Incident Response Lifecycle

Incident Response Lifecycle

flowchart LR
    A[Healthy Baseline] --> B[Incident Generated]
    B --> C[Telemetry Changes]
    C --> D[KQL Investigation]
    D --> E[Root Cause Identified]
    E --> F[Remediation]
    F --> G[Telemetry Recovery]
    G --> H[Alert / Query Recovery]
    H --> I[Incident Closed]
Loading

Skills Demonstrated

Azure Administration

  • Azure resource groups
  • Azure virtual machines
  • Azure Portal
  • Azure resource monitoring
  • Azure Activity Log

Monitoring

  • Azure Monitor
  • VM metrics
  • Log Analytics
  • Azure Monitor Agent
  • Data Collection Rules
  • Windows performance counters
  • Heartbeat monitoring

Alerting

  • Metric alerts
  • Scheduled query alerts
  • Alert rules
  • Action groups
  • Severity levels
  • Threshold configuration
  • Automatic resolution
  • Alert lifecycle investigation

Kusto Query Language

  • where
  • project
  • extend
  • summarize
  • arg_max
  • min
  • max
  • avg
  • count
  • iff
  • ago
  • bin
  • order by
  • top
  • render timechart
  • ingestion_time()

Windows Administration

  • Windows Server 2022
  • PowerShell
  • CIM
  • Windows Event Log
  • Performance counters
  • Process investigation
  • CPU utilization
  • Memory utilization
  • Storage capacity
  • VM Run Command

Troubleshooting

  • Incident triage
  • Baseline analysis
  • Performance troubleshooting
  • Agent troubleshooting
  • Availability troubleshooting
  • Event investigation
  • Capacity troubleshooting
  • Root-cause analysis
  • Evidence correlation
  • Recovery validation

Support Operations

  • Incident documentation
  • Evidence management
  • Alert correlation
  • Monitoring validation
  • Remediation planning
  • Incident closure
  • Git version control
  • GitHub documentation

Repository Structure

Azure-Monitor-Log-Analytics-Support-Lab/
│
├── Documentation/
│   ├── 01-Environment-Setup.md
│   ├── 02-Log-Analytics-Workspace.md
│   ├── 03-Azure-Monitor-Agent-DCR.md
│   ├── 04-VM-Metrics-Monitoring.md
│   ├── 05-Activity-Logs-Diagnostics.md
│   ├── 06-Azure-Monitor-Alerts.md
│   ├── 07-KQL-Log-Analysis.md
│   ├── 08-Incident-Investigation.md
│   └── 09-Root-Cause-Analysis.md
│
├── Help-Desk-Tickets/
│   ├── INC-001-High-CPU.md
│   ├── INC-002-Missing-Heartbeat.md
│   ├── INC-003-VM-Availability.md
│   ├── INC-004-Windows-Event.md
│   ├── INC-005-Disk-Space.md
│   ├── INC-006-Memory-Pressure.md
│   └── INC-007-Log-Query-Alert.md
│
├── KQL/
│   ├── Activity-Log-Queries.kql
│   ├── Disk-Space-Queries.kql
│   ├── Heartbeat-Queries.kql
│   ├── Incident-Investigation-Queries.kql
│   ├── Log-Query-Alert-Queries.kql
│   ├── Memory-Pressure-Queries.kql
│   ├── Missing-Heartbeat-Queries.kql
│   ├── Performance-Queries.kql
│   ├── VM-Availability-Queries.kql
│   ├── Windows-Event-Incident-Queries.kql
│   └── Windows-Event-Queries.kql
│
├── Screenshots/
│   └── 134 monitoring and incident screenshots
│
└── README.md

Portfolio Value

This project was built to demonstrate practical skills relevant to roles such as:

Azure Support Engineer
Cloud Support Specialist
Technical Support Engineer
IT Support Specialist
Systems Administrator
Cloud Administrator
Help Desk Technician
Infrastructure Support Technician
Microsoft Support Specialist
NOC / Monitoring Analyst

Rather than documenting only successful Azure resource creation, the project emphasizes:

something breaks
      ↓
support investigates
      ↓
evidence is gathered
      ↓
root cause is identified
      ↓
the issue is fixed
      ↓
recovery is proven

Interview Talking Points

What is Azure Monitor?

Azure Monitor provides monitoring and observability across Azure resources, applications, operating systems, metrics, logs, and alerts.

In this lab I used Azure Monitor to investigate:

CPU
VM availability
memory
storage
activity logs
agent health
alerts

What is Log Analytics?

Log Analytics provides a query environment for analyzing monitoring data stored in an Azure Log Analytics workspace.

I used KQL to investigate:

Heartbeat
Perf
Windows events
resource state
memory
disk capacity
incident timelines

What is Azure Monitor Agent?

Azure Monitor Agent collects guest operating-system monitoring data and forwards it according to Data Collection Rules.

In this environment it collected Windows telemetry from:

AZMON-WIN01

into:

LAW-AZMON-SUPPORT-LAB

What is a Data Collection Rule?

A Data Collection Rule determines what monitoring data should be collected and where it should be sent.

The lab used:

DCR-AZMON-WINDOWS

to collect Windows performance data.


How did you troubleshoot missing monitoring data?

I first determined whether the VM itself was unavailable or whether only the monitoring telemetry had stopped.

I used:

Heartbeat
VM state
Azure Monitor Agent status
Activity Log
recent telemetry timestamps

to isolate the problem.


How did you troubleshoot memory pressure?

I compared:

Windows physical memory
process working set
Memory\Available Bytes
Log Analytics Perf data
KQL timelines

The controlled worker consumed approximately 1 GB and the reduction was visible both locally and in Log Analytics.


How did you test alerting?

I built a scheduled KQL query rule that detected when:

Available memory < 2048 MB

I validated:

healthy state
alert condition false
controlled threshold violation
alert firing
remediation
telemetry recovery
condition clearing
automatic alert resolution

Project Outcome

The completed lab demonstrates the ability to:

Deploy monitoring
Collect telemetry
Query operational data
Build alerts
Simulate incidents
Investigate symptoms
Identify root causes
Remediate failures
Validate recovery
Document incidents

The final environment contains:

9 monitoring documentation files
7 completed support incidents
11 reusable KQL files
134 screenshots
Azure metric monitoring
Log Analytics monitoring
Azure Monitor Agent telemetry
Data Collection Rules
Azure Activity Log investigation
Metric alerting
KQL scheduled query alerting
Windows Server troubleshooting
End-to-end incident validation

Final Status

Azure Monitor Environment:
COMPLETE

Log Analytics Environment:
COMPLETE

Azure Monitor Agent:
VERIFIED

Data Collection Rule:
VERIFIED

Performance Monitoring:
VERIFIED

Heartbeat Monitoring:
VERIFIED

Activity Log Monitoring:
VERIFIED

Windows Event Monitoring:
VERIFIED

Metric Alerting:
VERIFIED

KQL Scheduled Query Alerting:
VERIFIED

Incident Investigations:
7 COMPLETE

Reusable KQL Library:
11 FILES

Screenshot Evidence:
134

Repository Status:
PORTFOLIO READY

About

Hands-on Azure Monitor + Log Analytics support lab featuring AMA/DCR, KQL, alerts, Windows telemetry, incident investigation, remediation, and recovery validation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors