Malmö University Publications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Detecting LLM-Based Web Honeypots Using Machine Learning
Malmö University, Faculty of Technology and Society (TS), Department of Computer Science and Media Technology (DVMT).
Malmö University, Faculty of Technology and Society (TS), Department of Computer Science and Media Technology (DVMT).
2026 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

One of the shortcomings of low and medium interaction honeypots is the static nature of their responses. To address this, solutions using Large Language Models (LLMs) have been proposed to equip them with dynamic response generation capabilities, serving to decrease detectability and consequently, improve attacker engagement. However, to ensure the efficacy of the systems, there is a need for rigorous evaluation of their detectability. While previous research has demonstrated machine learning classifiers capable of identifying static honeypots, this approach has yet to be applied to LLM-powered honeypots. Therefore, this thesis investigates the feasibility of utilizing machine learning technology to identify characteristics of LLM web honeypot responses for fingerprinting purposes. By applying novel evaluation techniques, we aim to provide developers and honeypot administrators with a tool for the assessment of honeypot detectability. To this end, a set of 1,597 URIs was collected and served as the basis for compiling a dataset. Legitimate responses, alongside generated responses from two open source honeypots (Galah, Beelzebub), were gathered and used to train a random forest classifier. The model proved capable of distinguishing between legitimate and honeypot responses, where response latency and lack of credible headers were shown to be of particular importance. We present a framework for fingerprinting LLM honeypots which, by delivering actionable results, could be used during both honeypot development and configuration.

Place, publisher, year, edition, pages
2026. , p. 37
Keywords [en]
Cybersecurity, Detection, Fingerprinting, Honeypots, Hypertext Transfer Protocol, Large Language Models, Machine Learning
National Category
Security, Privacy and Cryptography
Identifiers
URN: urn:nbn:se:mau:diva-85918OAI: oai:DiVA.org:mau-85918DiVA, id: diva2:2073066
Educational program
TS Systemutvecklare
Presentation
2026-06-02, NI:A0502, Nordenskiöldsgatan 1, Malmö, 11:30 (English)
Supervisors
Examiners
Available from: 2026-06-23 Created: 2026-06-16 Last updated: 2026-06-23Bibliographically approved

Open Access in DiVA

fulltext(1925 kB)61 downloads
File information
File name FULLTEXT02.pdfFile size 1925 kBChecksum SHA-512
ee030b36401b4f66f0e7a3f7cb0c03223361a64421b8bf44a3a1bab153dfaf29122e555aab1af718d05e621b3dfd58b6b0fc4ad904098ccdf638212af9c7a194
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Lundman, SamuelRosenquist, Johnny
By organisation
Department of Computer Science and Media Technology (DVMT)
Security, Privacy and Cryptography

Search outside of DiVA

GoogleGoogle Scholar
Total: 61 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 119 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf