🏞️QGISAdvanced⏱️ 4 mins read

Land Use / Land Cover (LULC) Classification

Published by GISTECHNEWS Editorial Team • Peer-Reviewed & Verified on QGIS 3.34+ LTR & Python 3.10+

Land Use and Land Cover (LULC) classification is the process of categorizing multi-spectral satellite imagery pixels into discrete environmental classes such as Water Bodies, Built-Up Infrastructure, Agricultural Cropland, Dense Forest, and Barren Land. While 'Land Cover' describes the physical surface material (e.g., deciduous trees), 'Land Use' describes socioeconomic human activities on that land (e.g., timber plantation or recreation). In this advanced tutorial, we deploy the Semi-Automatic Classification Plugin (SCP) and machine learning algorithms (Random Forest and Maximum Likelihood) in QGIS to execute an end-to-end classification workflow.

📋 Prerequisites

  • QGIS 3.34+ LTR with the Semi-Automatic Classification Plugin (SCP) installed.
  • A cloud-free multispectral satellite scene (Sentinel-2 L2A or Landsat 8-9 OLI).
  • Understanding of spectral reflectance curves and false color composites.

🛠️ Technical Environment

Required Software: QGIS Semi-Automatic Classification Plugin (SCP) (Recommended: QGIS 3.34 LTR)

Practice Dataset: Multispectral Sentinel-2 Surface Reflectance Stack

Source Portal: Copernicus Browser

CRS / Format: UTM (Multi-band GeoTIFF)

Step-by-Step Workflow & Methodological Execution

1

Module 1: Pre-Processing and Training Sample Collection

The quality of supervised classification depends almost entirely on the quality of your Training Input Data (Regions of Interest - ROIs): 1. In QGIS, open the SCP Dock and create a new Training Input file (`training_samples.scp`). 2. Set up Class IDs: Define Macroclass IDs (MC_ID): 1: Water, 2: Vegetation, 3: Built-Up, 4: Bare Soil. 3. Display False Color Composite: Render your image using NIR-Red-Green bands (e.g., Sentinel-2 8-4-3). Healthy vegetation displays as bright red, urban concrete as cyan/gray, and water as deep blue/black. This makes identifying pure endmembers intuitive. 4. Collect Pure Pixels: Use the ROI Polygon tool to draw 15-20 small, homogeneous polygons for each class across the entire image extent. Avoid mixed pixels at parcel edges!

2

Module 2: Evaluating Spectral Separability

Before running a classifier, you must verify that your training classes do not overlap spectrally: 1. Open the SCP Spectral Signature Plot tool and load your collected ROIs. 2. Analyze Band Overlaps: Water should display low values across all bands. Vegetation should show a massive peak in Band 8. Built-up surfaces should slope upward into the SWIR bands. 3. Compute Separability Metrics: SCP calculates the Jeffries-Matusita (JM) Distance and Transformed Divergence. A JM distance value between 1.9 and 2.0 indicates excellent statistical separability. If the JM distance between 'Agriculture' and 'Forest' is below 1.4, collect more distinct training samples or merge them into a single Macroclass.

3

Module 3: Classification Execution & Post-Processing

1. Select Classifier Algorithm: Choose Random Forest (robust against noise, non-parametric) or Maximum Likelihood. 2. Execute Classification: SCP generates a thematic raster where each pixel holds an integer code corresponding to your classes. 3. Sieve Filtering: Satellite classifications often suffer from the 'salt-and-pepper' effect (isolated single-pixel noise). Navigate to Raster -> Analysis -> Sieve to remove pixel clusters smaller than 10 cells (1,000 m² for Sentinel-2), replacing them with the dominant neighboring class. 4. Vectorization: Convert the cleaned thematic raster into polygon vectors via Raster -> Conversion -> Polygonize (Raster to Vector).

⚠️ Common Errors & Troubleshooting

❌ Classifier misclassifies urban asphalt as water

💡 Resolution: Asphalt and deep water can have similar low NIR reflectance; add SWIR bands or NDWI indices into the input stack to separate them.

❌ SCP plugin crashes during model training

💡 Resolution: Ensure you have at least 50–100 training pixels per class and that no ROI polygons fall outside the raster boundary.

💡 Expert Tips & Best Practices

  • Use a stratified random sample design to collect independent validation polygons for honest confusion matrix accuracy assessment.
  • Always filter out cloud contaminated pixels prior to training machine learning classifiers.

🐍 Python Scikit-Learn Random Forest Satellite Classifier

import rasterio
import numpy as np
from sklearn.ensemble import RandomForestClassifier

# Open 4-band satellite image (Blue, Green, Red, NIR)
with rasterio.open("multispectral_scene.tif") as src:
    img = src.read() # Shape: (4, rows, cols)
    profile = src.profile

# Open rasterized training mask (1: Water, 2: Veg, 3: Urban)
with rasterio.open("training_labels.tif") as lbl_ds:
    labels = lbl_ds.read(1)

# Reshape image array from (Bands, Rows, Cols) to (Pixels, Bands)
bands, rows, cols = img.shape
X = img.reshape(bands, rows * cols).T
y = labels.flatten()

# Filter labeled training pixels (exclude 0 / background)
mask = y > 0
X_train = X[mask]
y_train = y[mask]

# Train Random Forest Classifier
clf = RandomForestClassifier(n_estimators=100, n_jobs=-1, random_state=42)
clf.fit(X_train, y_train)

# Predict entire satellite scene
predicted_pixels = clf.predict(X)
lulc_map = predicted_pixels.reshape(rows, cols)

# Save classified LULC raster
profile.update(dtype=rasterio.uint8, count=1, nodata=0)
with rasterio.open("lulc_random_forest.tif", "w", **profile) as dst:
    dst.write(lulc_map.astype(rasterio.uint8), 1)

print("Supervised Random Forest classification finished.")

🔗 Related Tutorials & Practical Workflows