Intelligent Web Attack Classification Using Ensemble Machine Learning Models

—Secure injection and traversal based attacks on web
Attacks on applications. is a growing issue, as attackers are
becoming more and more familiar with the attacks. Armed with
increasingly sophisticated attacks, payloads that can execute even
more dangerous tasks. Bypass traditional rules-based defence.
In this paper we present a multi-class HTTP request classifier
built on The features include character-level TF-IDF features and
LightGBM ensemble. approach of extended with the SHAP-based
post-hoc explainabil- ity to support. Analysis of Model Decisions
at Analyst level. Following experiments were carried out using the
datasets “CSIC 2010” and ”ECML/PKDD 2007”, which cover:
The accuracy of LightGBM In 11,347 labeled samples in five
categories, is shown. 94.32higher in value than the lowest value. A
controlled ablation is carried To test the accuracy of the character
n-grams over the In contrast, range-based tokenization is 6.12
points better than A word-based tokenization with a vocabulary
of [2,4] and 10,000 words. at the “knee” of the accuracy-memory
curve. The trained pipeline can be fitted in the memory of 80
MB and can run on a normal A CPU chipset that doesn’t
support GPUs, for use with CPU- only hardware. ModSecurity
compatible WAF deployment. In In the SHAP attribution maps,
the SQL metacharacters are shown as well as script-injection.
The following are the key factors: For their class, they were given
tokens, and sequences that lead to path traversal. An operator
of a WAF removes audit trail for every blocked request.