IUI A S A R Y K U N I V E R S I T Y FACULTY OF INFORMATICS Generator of Vulnerable Web Applications Bachelor's Thesis MAREK GELETA Brno, Spring 2025 IUIA S A R Y K U N I V E R S I T Y FACULTY OF INFORMATICS Generator of Vulnerable Web Applications Bachelor's Thesis MAREK GELETA Advisor: doc. RNDr. Jan Vykopal, Ph.D. Department of Computer Systems and Communications Brno, Spring 2025 Declaration Hereby I declare that this paper is m y original authorial work, which I have worked out o n m y o w n . A l l sources, references, a n d literature used or excerpted during elaboration of this work are properly cited and listed i n complete reference to the due source. Marek Geleta Advisor: doc. R N D r . Jan Vykopal, P h . D . iii Acknowledgements I w o u l d like to express m y sincere gratitude to m y thesis consultant, doc. R N D r . Jan Vykopal, Ph.D., for his invaluable guidance, support, and expertise throughout the course of this thesis. Funded by the European U n i o n under Grant Agreement number 101087529. V i e w s and opinions expressed are however those of the author (s) only and do not necessarily reflect those of the European Union or European Research Executive Agency. Neither the European U n i o n nor the granting authority can be held responsible for them. Use of AI Tools During the preparation of this thesis, the following A I tools were utilized: • Grammarly: For grammar checking and correcting language mistakes. • GitHub Copilot: To accelerate code development and provide code suggestions. • ChatGPT: For i m p r o v i n g writing style, refining explanations, and suggesting structural improvements to the text. I declare that I used these tools i n accordance with the principles of academic integrity. I checked the content and took full responsibility for it. iv Abstract Capture-the-Flag competitions frequently employ the attack-defence format, but bringing that format into university courses is hindered by the labor of creating fresh vulnerable services—re-using past targets encourages solution sharing between students. This thesis introduces a generator that, from a brief text specification selecting vulnerability types and difficulty levels f r o m a predefined catalogue, builds a containerised service ready for deployment into the F A U S T attackdefense framework. The core business logic is fixed, yet the generator randomises A P I identifiers, interface theme, and the exact placement of each vulnerability, yielding a distinct exploit path for every requested flaw. It produces a runnable container image together w i t h FAUST-compatible checker scripts, enabling instructors to publish a fully functioning, exploitable target without manual adjustment. A p i lot classroom trial confirmed that the tool generated applications that deployed out of the box and behaved as intended, showing that automated diversification can make attack-defence C T F exercises practical for teaching while significantly cutting preparation effort. Keywords vulnerability generator, web application security, cybersecurity education, Capture The Flag, S Q L injection, authorization flaws, FAUST Attack-Defence v Contents 1 Introduction 1 2 Background 4 2.1 C T F Formats 4 2.2 Related Work and Existing Solutions 5 2.3 FAUST C T F Framework 6 2.4 Technologies Used 10 2.5 Software Security Vulnerabilities 13 3 Design 18 3.1 Requirements 18 3.2 Generator design 21 4 Implementation 24 4.1 Generator core 24 4.2 Configuration generator 31 4.3 Baseline code template 32 4.4 Validation scripts 36 4.5 Automation 38 5 Evaluation and Testing 44 5.1 Testing Procedure 44 5.2 Observations 45 5.3 Improvements 47 6 Conclusions and Future Work 48 A Content of the attachment 51 Bibliography 52 vi List of Figures 2.1 Scoreboard from FAUST C T F 2024 [14] 8 2.2 A FAUST C T F architecture diagram illustrating interaction between the architecture components and players during a game 11 3.1 High-level architecture diagram 21 4.1 User Interface Screenshots of the Application 34 4.2 C I / C D Job continuing into validation phase 41 vii 1 Introduction The landscape of digital threats continues to grow and evolve, w i t h cybersecurity incidents affecting organizations of all sizes across all sectors. A s digital infrastructure becomes increasingly more important to m o d e r n society, the sophistication and frequency of cyberattacks have grown i n parallel. According to the European Union Agency for Cybersecurity ( E N I S A ) , the number of reported data compromises increased by 78% i n 2023 compared to 2022, w i t h this u p w a r d trend continuing into 2024. E N I S A also observed a significant rise i n the exploitation of vulnerabilities and reported that the impact of system and human errors more than tripled i n 2023 [1]. This rapidly evolving threat landscape demands skilled cybersecurity professionals capable of defending against increasingly complex attacks. Education i n cybersecurity presents unique challenges that traditional pedagogical approaches struggle to address. While theoretical knowledge forms an essential foundation, practical skills development requires hands-on experience w i t h real-world scenarios. Gamification has emerged as a particularly effective educational strategy i n this domain. Capture The Flag (CTF) competitions provide dynamic, immersive, and challenging learning environments. Participants actively apply security concepts to identify and exploit vulnerabilities in controlled settings. These competitions transform abstract security principles into concrete problems, targeting the improvement of both technical skills and analytical thinking. Universities, corporations, and cybersecurity communities worldwide have embraced C T F competitions as valuable educational tools. [2, 3] Academic institutions incorporate them into curricula to complement theoretical coursework, w h i l e companies utilize them for employee training and recruitment. Events such as D E F C O N C T F [4] and picoCTF [5] draw thousands of participants each year, building strong and engaging learning ecosystems. A m o n g C T F formats, Attack-Defense ( A / D ) competitions stand out for their comprehensive approach to security education. Unlike jeopardy-style C T F s that focus primarily on exploitation, A / D competitions require participants to simultaneously defend their systems while attacking others'. This environment develops professional skills 1 i . I N T R O D U C T I O N often overlooked i n other C T F formats: vulnerability patching, i n trusion detection, teamwork, and strategic prioritization under time constraints, where participants must make real-time decisions about resource allocation between offensive and defensive operations. Educational institutions face a significant challenge when incorporating C T F exercises into recurring courses: maintaining assessment integrity across terms. Reusing identical challenges enables students to share solutions from previous iterations, undermining the educational value of the exercise. Conversely, developing entirely new scenarios for each course iteration requires substantial resources and expertise. Instructors must balance the need for novel challenges against practical constraints of time and expertise, particularly w h e n creating complex A / D scenarios that require multiple interconnected vulnerabilities and supporting infrastructure. This thesis addresses these challenges through the development of a specialized generator for vulnerable web applications. The generator produces applications w i t h consistent core functionality but randomized implementation details: different vulnerability types and locations, varied visual appearances, and randomized A P I endpoints. This approach preserves the pedagogical objectives while presenting students w i t h unique technical challenges each time. The generator supports integrating S Q L injection and authorization flaws, w i t h configurable difficulty levels to accommodate different learning stages and educational objectives. The generated applications are fully compatible w i t h the FAUST Attack-Defense framework [6], an established platform for hosting A / D competitions. Each application includes essential components for deployment: containerized services, database initialization scripts, and checker scripts for automated functionality verification and flag placement. This integration significantly reduces the overhead associated with creating and deploying new A / D scenarios, enabling instructors to focus on educational aspects rather than technical implementation details. The remainder of this thesis is organized as follows: Chapter 2 provides background information on C T F competitions, describes the FAUST Attack-Defense framework, reviews existing vulnerability generation tools, and covers essential technologies and vulnerabilities addressed i n this work. Chapter 3 presents the design requirements 2 i . I N T R O D U C T I O N and architectural decisions that guided the development of the generator. Chapter 4 details the implementation of key components, the core structure of the generator, validation scripts, and the intended vulnerability details. Chapter 5 evaluates the generator through testing i n a real educational environment and describes improvements made following user feedback. Finally, Chapter 6 summarizes the contributions of this work and outlines directions for future research. 3 2 Background To provide context, the chapter first introduces Capture The Flag competitions and their most c o m m o n formats, then surveys current automated vulnerability generation tools and their limitations, introduces the FAUST Attack-Defense architecture, and closes w i t h a brief overview of the key technologies and security flaws used i n the later implementation. 2.1 CTF Formats Kucek et al. defines Capture the flag (CTF) as "a computer security competition with challenging exercises from different categories (e.g., network security, application security, mobile security, etc.). Typically, the goal of this game is to find a flag, i.e., the solution to the given problem (which is often solved through multiple steps or tasks within an exercise). A flag can be, for example, a string i n the format f lag{sometext} but also an image or another element defined by the CTF game designers." [7] The two most used formats i n competitions are jeopardy-style and attack-defense ( A / D ) . Jeopardy-Style CTFs Jeopardy-style C T F s are organized around independent challenges categorized by topic (e.g., cryptography, web, reverse engineering). Each challenge, once solved, reveals a unique flag that can be submitted for points. [8] A n example of a web-based Jeopardy-style C T F challenge involves a web application with an administrator panel that is vulnerable to unauthorized access. Participants are tasked w i t h identifying and exploiting a security flaw that allows them to bypass authentication mechanisms and access the administrator panel, where the flag is located. Attack-Defense CTFs In A / D CTFs, each team runs its o w n instance of intentionally vulnerable services (typically inside a virtual machine, or "Vulnbox"). Teams 4 2. B A C K G R O U N D must defend these services f r o m exploitation while simultaneously attacking the services of other teams. Points are awarded for capturing flags from other teams, protecting your own, and maintaining service availability ( S L A ) . Because this format forces participants to handle realistic conditions - patching, monitoring, writing exploits, maintaining uptime-it is widely used i n education and team training. 2.2 Related Work and Existing Solutions Several tools are available to automate the generation of vulnerable applications and environments for cybersecurity training. They aim to facilitate the creation of realistic attack scenarios for education and assessment. However, each exhibits limitations i n terms of flexibility, m o d e r n technology adoption, and integration w i t h Attack-Defense ( A / D ) competition frameworks. SiteGenerator One of the earliest attempts is the O W A S P SiteGenerator project [9]. SiteGenerator was designed to dynamically generate websites containing a configurable set of vulnerabilities for benchmarking security scanners and training purposes. However, the project remains incomplete and has not been actively maintained since 2007. Additionally, its generated applications are primarily based on outdated technologies such as ASP.NET and PHP5. Vulnerability integration i n SiteGenerator operates on a binary on-off basis - developers must explicitly include both vulnerable and secure code paths, letting users choose which version to deploy. This limits both the granularity of difficulty configuration and the variety of scenarios that can be generated automatically. V W G e n V W G e n [10] shares similar limitations. A l t h o u g h it automates the generation of web applications w i t h vulnerabilities like SQL injection and local file inclusion, it is an experimental tool without active development since 2017 that primarily targets older web stacks and requires manual definition of vulnerable code snippets. V W G e n lacks support for m o d e r n web development frameworks and does 5 2. B A C K G R O U N D not provide fine-grained control over vulnerability difficulty or placement. Furthermore, neither SiteGenerator nor V W G e n allows for easy integration w i t h modern A / D competition platforms, requiring substantial manual adaptation to incorporate their outputs into structured competitions. SecGen In contrast, SecGen [11] approaches the problem f r o m a higher abstraction level by generating entire vulnerable virtual machine environments rather than individual applications. SecGen is primarily used to create randomized security scenarios for Capture The Flag (CTF) events and academic exercises. W h i l e powerful for infrastructure-level challenge generation, it does not focus on finegrained control over web application vulnerabilities and lacks direct compatibility w i t h established A / D platforms like the FAUST framework. Generated scenarios require additional effort to integrate into competition environments and lack modularity at the application level. This thesis addresses these limitations by introducing a novel generator for vulnerable web applications built on modern technologies, specifically Node.js and Express. Unlike previous tools, the proposed generator allows granular specification of vulnerability types and difficulty levels, enabling educators to tailor scenarios to their learning objectives. Most importantly, the generated applications are fully compatible w i t h the F A U S T Attack-Defense framework, facilitating seamless deployment i n A / D C T F competitions without requiring additional adaptation or manual integration efforts. 2.3 FAUST CTF Framework The FAUST C T F framework [12] is an open-source system developed by the F A U Security Team [6], designed to facilitate attack-defense ( A / D ) style Capture The Flag competitions. Its architecture comprises of several components that cooperate i n order to manage services, monitor their availability, handle flag placement and retrieval, and calculate scores as shown i n Figure 2.2. 6 2. B A C K G R O U N D Vulnbox Each participating team is p r o v i d e d w i t h a virtual machine ( V M ) k n o w n as a vulnbox. This V M hosts a set of intentionally vulnerable services that are identical across all teams. Teams are responsible for defending their o w n Vulnboxes while attempting to exploit v u l nerabilities i n other teams' Vulnboxes to capture flags [12]. Usually, teams have superuser access to the Vulnbox, allowing for more efficient administration a n d creative ways of patching vulnerabilities and detecting vulnerabilities being exploited - for example, running a packet capture. Services Services are applications r u n n i n g o n the Vulnbox, deliberately designed with security flaws to be exploited during the competition. The services are commonly implemented i n the form of Docker containers running o n the vulnbox. Each service is accompanied b y a checker script that verifies its functionality and availability [12,13]. Checker Scripts (Checkerbots) For each service, a corresponding checker script—often referred to as a checkerbot—is implemented in Python or Go. These scripts perform flag placement, service functionality validation, and flag retrieval. Checker scripts utilize the provided c h e c k e r l i b library, which simplifies their development and ensures compatibility w i t h the scoring system [13]. Gameserver (Controller) The Gameserver, also k n o w n as the controller, orchestrates the competition by managing ticks, generating n e w flags, executing checker scripts, accepting flag submissions, calculating scores, and hosting a scoreboard. [12, 6]. Throughout the game, players can see the scores in real time o n a scoreboard as show i n Figure 2.1 7 2. B A C K G R O U N D FAUST CTF 2024 Privacy Legal Scoreboard /Cy^N TeamGeramy A 5271.06 ft * -22.00 7697.43 0 2494.15 K-58.4 SecretChannel service Ivm asm chat ft 7453.12 #5713.60 4*0.00 ft S0.00 SO.00 SO.OO 1425.51 ©2416.21 0 2010.91 0 S up up 2478.56 -195.42 Offense Defense SLA Total 38125.60 -347.05 17490.25 55268.79 ft 5673.03 ft ft 3901.25 S. -66.67 9734.28 ^-74.99 ©1816.06 *-52.51 ©2268.12 ft 8849.11 ft 0.00 ft ^-66.61 ^0.00 2419.35 02112.24 © * recovering 2478.56 -373.05 LP © 2174.59 ft 423.67 ft S. -13.28 7834.52 0 *-177.71 1769.29 © recovering 1535.46 38835 19 -824 81 16492.59 54502.96 3. Superflat 0 1784.88 W-413.03 ©1519. up © up 1769.29 249415 -290.03 0.23 14918.15 50606.33 ft 5173.41 0 2244.74 ft 4796.46 317.29 ©2205.77 ft9575.92 ftO.OO ft273.13 ft 0.00 01816.06 0 -378.49 0 Figure 2.1: Scoreboard from FAUST C T F 2024 [14] Ticks The competition is segmented into intervals called ticks. These intervals are usually i n the range of 30 seconds to 5 minutes, set by the competition administrators. Each tick, the gameserver re-runs the checker scripts and updates the scoring. [12]. Flags Flags are r a n d o m strings following a predetermined format (e.g., FAUST_ [32 random c h a r a c t e r s ] ) . They are central to the Attack-Defense C T F format. Flags serve as concrete goals for attackers and defenders, creating a measurable competitive element. In the FAUST framework [12], flags follow a specific lifecycle: Flag Generation and Placement, Flag Stealing, and Flag Defense. A l l of the mentioned can be seen i n Figure 2.2 except for Flag Defense, w h i c h w o u l d be hard to illustrate as successful flag defense results in other teams' inability to r u n an exploit successfully. 8 2. B A C K G R O U N D Flag IDs Each flag is associated with a unique identifier known as a flag ID. This identifier references specific instances of data where flags are stored, for example, an identifier of a user who stored the flag in their private notes. The flag IDs facilitate interactions between the gameserver and participants, important for both service functionality validation and exploit development [13]. Flag IDs are public to all the teams as some vulnerabilities may not be exploitable without their knowledge. O n the contrary a service may not provide flag IDs if their authors decide so. This is the case w h e n the service's vulnerabilities are exploitable without arbitrary information. A n example of this is an S Q L injection vulnerability that allows users to gather all flags f r o m the database without any additional information. Flag Generation and Placement At the beginning of each tick, the Gameserver generates unique flags for each team and service. The checkerbot script then inserts these flags into each team's service instance, storing them in predefined locations such as database records, text files, or application data structures. After the checkerbot inserts a flag into the service's data store, it saves the corresponding flag ID so it can be accessed by all the teams. The insertion process usually involves: 1. Creating a regular user account i n the service. 2. Storing the flag as private information (e.g., in a private message or note). 3. Recording the flag ID i n the gameserver database. Older flags are automatically invalidated after a predefined n u m ber of ticks, creating a rolling w i n d o w of valid flags. Flag Stealing Teams exploit vulnerabilities i n opponents' services to extract flags placed by the checkerbot. Every team is p r o v i d e d w i t h flag IDs of other teams as for some vulnerabilities The Flag IDs are important 9 2. B A C K G R O U N D w h e n crafting exploits. They provide a direct reference to specific data entries w i t h i n the application. For instance, if a service has an authorization flaw that allows access to any user's data given their user ID, k n o w i n g the flag ID, w h i c h corresponds to a user ID, enables an attacker to retrieve the flag without proper authorization. A successful attack involves: 1. Identifying and exploiting a vulnerability i n the target service. 2. Using the exploit to access the protected flag data. 3. Extracting the flag from the service. 4. Submitting the stolen flag to the Gameserver. The Gameserver validates submitted flags and awards attack points if the flag is valid and was stolen from another team's service. Flag Defense Teams must also defend their o w n services by patching vulnerabilities to prevent flag theft. The checkerbot periodically verifies that: 1. The service is accessible and responding correctly. 2. Previously placed flags can still be retrieved using legitimate access methods. 3. N e w flags can be placed successfully. A team earns defense points w h e n their flags remain secure and their service maintains functionality. Points are deducted w h e n flags are successfully stolen or w h e n the service becomes unavailable or dysfunctional. 2.4 Technologies Used This section presents an overview of the technologies applied i n the development of the individual parts of the vulnerability generator described i n Chapter 4. Each section also provides a brief explanation of the purpose each technology serves w i t h i n the generator or the resulting vulnerable application. 10 2. B A C K G R O U N D Service A Service B Team 1 Player Functionality check + place Mags Checkerbot A Checkerbot B I |Score| Flag IDs Service A Service B scoreboard Flag ID endpoint inFlag submission service |1. Gather flag ID for service 3. Submit flag |2. Exploit and gather flag| Team 2 Player Figure 2.2: A FAUST C T F architecture diagram illustrating interaction between the architecture components and players during a game. Core Libraries Jinja2 Jinja2 is a templating engine that supports embedding control structures, variable substitutions, and expressions w i t h i n text-based templates. It is primarily used in conjunction with the Python programming language. Template engines of this kind facilitate the automated generation of structured text, such as H T M L or configuration files, based on predefined templates and variable input [15]. In the presented vulnerability generator, Jinja2 is used to generate randomised source code of the vulnerable application from a template. SQLGlot S Q L G l o t is a software library for parsing, transforming, and generating S Q L queries. It converts S Q L statements into abstract syntax trees (ASTs), w h i c h can be programmatically modified and re-rendered i n various S Q L dialects. This functionality is relevant i n contexts where applications need to interact w i t h multiple database systems using different dialects of S Q L [16]. This library is used to allow support for multiple S Q L dialects i n the generated vulnerable application, allowing for more variability in the generated application. Express.j s Express.js is a web application framework for the JavaScript programming language, designed to operate on the Node.js [17] run- 11 2. B A C K G R O U N D time environment. It provides a set of features for building web servers and APIs, including mechanisms for h a n d l i n g H T T P requests and responses, defining routing logic, and managing middleware functions. Expresses follows a m o d u l a r approach, allowing developers to integrate third-party libraries and middlewares to extend its core functionality. W h i l e it is often used i n the development of RESTful APIs and web applications, it does not impose specific architectural patterns, leaving design decisions to the developer. This flexibility makes it suitable for a range of applications, from simple web servers to more complex backend services [18]. The generated vulnerable application uses Express.] s i n its source code to handle H T T P requests, define routing, and manage middlewares for server-side functionality. Data Formats Y A M L Y A M L Ain't M a r k u p Language is a human-readable data serialization standard designed for configuration files and data exchange. It supports the representation of hierarchical data structures and is commonly used to define configuration settings i n a readable format. Y A M L files are often employed i n automation pipelines and software configuration tasks [19]. The generator uses Y A M L configuration files for specifying details about the vulnerabilities that can be inserted into the resulting application. Development and Testing Tools SQLMap S Q L M a p is an open-source software tool used to automate the detection and exploitation of S Q L injection vulnerabilities. It supports a range of techniques for identifying and verifying S Q L injection flaws, and it can interact w i t h various relational database management systems. S Q L M a p is used both i n security testing and educational environments to demonstrate the impact of S Q L injection vulnerabilities. One of its advanced features is the usage of tamper scripts that can tamper the S Q L injection payload i n order to evade an input filter or sanitization [20]. S Q L M a p is used i n this work as a means to automatically validate the presence of S Q L i vulnerabilities. 12 2. B A C K G R O U N D Infrastructure Components Docker Docker is a platform for containerization that allows software and its dependencies to be packaged into isolated environments called containers. Containers provide a consistent execution environment across different systems by encapsulating application code and dependencies while sharing the host operating system kernel. This approach is widely used to improve reproducibility and isolation i n software deployment [21]. The generated vulnerable application is provided w i t h a Docker configuration that allows it to be containerised and deployed on the Vulnbox. GitLab C I / C D G i t L a b C I / C D is a remote system for automating software build, test, and deployment processes using declarative configuration files. It supports the definition of pipelines, execution of jobs in defined stages, and integration w i t h containerized environments. This tool is used to implement continuous integration and deployment practices in software development workflows [22]. GitLab C I / C D simplifies the use of the generator by running it in a remote environment, eliminating the need for local setup by the user. 2.5 Software Security Vulnerabilities " A 'weakness' is a condition in a software, firmware, hardware, or service component that, under certain circumstances, could contribute to the introduction of vulnerabilities." [23] A software vulnerability is an instance of one or more weaknesses i n the design or implementation of software that could be exploited to compromise the confidentiality, integrity, or availability of a system. [24] These vulnerabilities arise from errors i n code, design decisions, configurations, or operational practices that create opportunities for unauthorized access or manipulation of software components. The C o m m o n Weakness Enumeration (CWE) provides a standardized classification system for software and hardware weaknesses. This framework enables the categorization, identification, and mitigation of vulnerabilities across different platforms and technologies [25]. 13 2. B A C K G R O U N D Assigning C W E identifiers to vulnerability types allows for more precise identification and communication about the core issue a n d enables researchers and developers to observe trends in software weaknesses over time. This categorization facilitates the development of targeted mitigation strategies and supports the creation of comprehensive security standards. For instance, the projects like O W A S P Top Ten, or C W E Top 25 utilize C W E classifications to group vulnerabilities by their underlying causes, rather their specific technical details, prov i d i n g a more logical framework for identification and remediation guidance [26,27]. Unexploitable Vulnerabilities Not all code weaknesses that appear to be vulnerabilities can be meaningfully exploited i n practice. These "unexploitable vulnerabilities" may exist due to environmental constraints, compensating controls, or implementation details that prevent an attacker from constructing a viable attack path. In educational contexts, unexploitable vulnerabilities create particular challenges as they can frustrate learners who correctly identify a theoretical weakness but cannot demonstrate its impact. For vulnerability generators intended for educational use, it is important to ensure that the introduced flaws are exploitable - students should be able to confirm their findings through successful exploitation. SQL Injection SQL Injection (SQLi) vulnerabilities occur when untrusted user input is improperly handled by a web application and directly incorporated into S Q L queries without adequate sanitization or validation. This allows attackers to alter the intended execution of S Q L commands, leading to unauthorized access, data leakage, or even data destruction [28]. The CWE-89 "Improper Neutralization of Special Elements used in an S Q L C o m m a n d " [29] is commonly used w i t h S Q L i vulnerabili- ties. Example of SQL Injection i n Python def login(request): username = request.get("username") 14 2. B A C K G R O U N D password = request.get("password") query = f"""SELECT * FROM users WHERE username = '{username} 1 AND password = '{password}';""" database.execute(query) Attacker's Input: • username: admin • password: ' OR = Resulting S Q L Query: SELECT * FROM users WHERE username = 'admin 1 AND password = '' OR ' 1 ' = ' 1 ' ; The injected condition OR ' 1' =' 1' always evaluates to TRUE, allowing the attacker to bypass authentication and gain unauthorized access. Second-order SQL Injection Second-order S Q L injection occurs w h e n malicious input is stored by the application and later used i n a S Q L query without proper revalidation or sanitization. Unlike traditional first-order S Q L i , where the injection is immediately executed, second-order S Q L i exploits rely on delayed execution, often triggered d u r i n g a subsequent action or request. This makes detection more challenging, as the initial input may appear harmless. For example, an attacker may insert a payload into a user profile field, w h i c h is later used i n a vulnerable backend query during login or reporting processes. Relevance of SQL Injection in Cybersecurity Education While the prevalence of S Q L i vulnerabilities has declined i n modern applications due to the widespread adoption of Object-Relational M a p ping ( O R M ) frameworks a n d parameterized queries, S Q L i remains 15 2. B A C K G R O U N D a critical component i n cyber security education. Teaching S Q L i provides foundational insights into input validation and filter bypasses, query manipulation, and the broader category of injection attacks. Injection attacks, including SQLi, remain highly relevant, as they continue to appear in the O W A S P Top Ten Web Application Security Risks, highlighting the persistent threat posed by improperly handled user input [27]. Authorization Flaws Authorization flaws occur w h e n a system fails to properly enforce access controls, allowing users to perform actions or access resources that should be restricted. After examining various types of authorization vulnerabilities i n the C W E framework, we have chosen to focus on CWE-863 "Incorrect Authorization" [30] for our implementation. While the We selected CWE-863 as our primary focus because it represents a realistic scenario i n real-world applications: developers typically attempt to implement authorization controls but fail to do so correctly. According to the C W E definition, this occurs w h e n "the product performs an authorization check w h e n an actor attempts to access a resource or perform an action, but it does not correctly perform the check" [30]. Unlike completely missing authorization checks ( C W E - 862 [31]), applications w i t h CWE-863 flaws implement some form of access control logic that contains subtle implementation errors, such as failing to verify permissions in all execution paths, applying checks inconsistently across different functions, or improperly handling edge cases. These flaws are particularly insidious as they may not be apparent during normal application usage, but can be abused by attackers w h o can exploit the flawed logic. Relevance in Cybersecurity Education Authorization vulnerabilities continue to be widespread i n m o d e r n applications, with "Broken Access Control" ranking first in the O W A S P Top Ten Web Application Security Risks since 2021 [27]. These flaws often arise from complex business logic and can be challenging to identify through automated testing alone. Teaching proper autho- 16 2. B A C K G R O U N D rization control implementation prepares students to design secure applications where every access operation is properly verified. 17 3 Design This chapter first states the generator's functional and non-functional requirements and then explains its architecture at two levels: the overall design and the structure of each component. 3.1 Requirements The generator is meant to be mainly used by academic staff and C T F administrators when preparing a C T F competition or assignment; hence, we expect the user to have an essential experience and understanding of the technologies used. Objectives The primary objective of the design process was to develop an architecture that satisfies the following requirements. Easy to use For the tool to be easy to use, it has to require m i n i m a l input from the user while providing the m a x i m u m outcome possible using that input. The input is expected to be i n human-readable text format, such as "l:easy:sqli" - meaning the user wants to generate an application w i t h one easily exploitable S Q L injection vulnerability. It should also allow using more advanced features for in-depth customization. This feature should be strictly optional, as the tool has to be able to function only w i t h m i n i m a l configuration from the user. Platform independent Platform independence refers to a tool's ability to operate seamlessly across various m o d e r n C P U architectures and operating systems without needing any source code alterations. One effective strategy to achieve this is using containerization technologies such as Docker [21] and P o d m a n [32]. These solutions not only simplify the development and deployment but also enhance user experience, as the containerized environment eliminates the necessity for users to install additional software on their local machines. Furthermore, by hosting the containerized tool i n an external remote environment (configurable Gitlab C I / C D ) , users can engage 18 3. D E S I G N w i t h an online service dedicated to executing the tool. This approach dramatically improves the platform independence, allowing users to access and utilize the tool f r o m any device, while also enhancing the overall usability of the system. The combination of these technologies ensures a flexible, platform-independent, and user-friendly experience. Capable of generating fully functional and intentionally vulnerable web applications The p r i m a r y function of the tool is to generate web applications that contain specific security vulnerabilities. These vulnerabilities must be deliberately embedded in a controlled manner. In addition, the generator must include a validation mechanism to ensure that the created applications function correctly, contain the intended vulnerabilities, and allow for their successful exploitation. This involves generating test scripts that check for vulnerabilities and confirm the application's expected behaviour. The validation process should be reliable, automated, and capable of detecting both functional errors and missing or unexploitable vulnerabilities i n the expected locations. The generated application should also mimic realworld scenarios to provide a realistic training environment. Compatibility with FAUST A D Service Format The generated application should be a plug-and-play F A U S T A / D service that does not require any additional code or configuration apart from the expected A / D C T F setup configuration overhead. This means that the generated service has to be deployable using Docker and contain a checkerbot [13] responsible for healthchecks and flag distribution. Reusable and r a n d o m i z e d The m a i n purpose of the tool is to be utilized i n an educational setting for exams, student competitions, or other types of cybersecurity training. To ensure fairness i n these exercises, it is important that the assignments vary each term. The tool operator should be able to generate a relatively large number of applications that differ i n vulnerabilities, database schema, code structures, variable names, and A P I endpoint names. B y d o i n g so, the tool can produce unique but functionally similar applications for 19 3. D E S I G N different use cases, preventing students or competitors from relying on predefined solutions. Functional Requirements 1. The tool must generate web applications vulnerable to: • S Q L Injection • Authorization flaws 2. The tool must support configuration parameters that define: • Type and number of vulnerabilities • Complexity or security level of the application 3. The generated applications must be containerized using Docker or Podman. 4. The generator must randomize: • Variable names and values • Database schemas and S Q L dialect used • Endpoint paths 5. The tool must provide a validation script that: • Confirms that vulnerabilities are exploitable • Ensures core application features are working 6. The generator must support exporting the generated application i n a format compatible w i t h FAUST A D Service. Non-functional Requirements 1. The generator should be operable on major operating systems (Linux, Windows, macOS) and C P U architectures (x64, arm64). 2. The tool should be extensible, allowing the addition of n e w vulnerability types or templates i n the future. 20 3. D E S I G N vulnerability definitions User input Application baseline Configuration generator config.yaml Code template Generator core Figure 3.1: High-level architecture diagram 3. Generation and validation should complete i n a reasonable time (less than 5 minutes for basic configurations). 4. Documentation and usage instructions must be provided. 5. The generated application should be lightweight and not require external proprietary software or services. 3.2 Generator design The design we chose satisfies all requirements and consists of three main interconnected parts (shown i n Figure 3.1) that work i n a chain to produce the requested vulnerable web application. Generator core The generator core, written i n Python, serves as the central component responsible for processing a configuration file i n a predefined Y A M L format. This configuration defines key parameters of the requested service, including the database management system to be used, database table names, frontend theme, and specific vulnerability characteristics. These parameters are used i n combination w i t h a baseline code template for the vulnerable application. Based on this input, the generator core assembles the final output. A s part of this process, it transforms the requested S Q L queries into intentionally vulnerable forms. B y default, the generated queries are i n the form of prepared statements [33] that are not vulnerable to S Q L i . Trans- 21 3. D E S I G N formation into the vulnerable form first includes transpilation of the query to the selected S Q L dialect. After the query is transpiled, the S Q L i vulnerability is introduced by using direct string concatenation of user-supplied variables into the S Q L query instead of the prepared statement. The generator core also outputs a FAUST checkerbot script and validation scripts to test the functionality and vulnerability of the generated application. Application Baseline The application baseline contains a code template of the vulnerable application w i t h specific spots pre-marked where vulnerabilities or specific code blocks can be inserted. This way, the generator core can modify the template dynamically based on the provided configuration. Another part of the baseline is a configuration file w i t h vulnerability definitions. The codebase handles the following functionality: user authentication, authorized data operations, private text notes creation, sharing, and retrieval. Registered users can view a list of all their notes and their content. They can also use the sharing feature to make the selected notes accessible to other users of their choice. This functionality was chosen as it is simple for the student to understand in a relatively short period of time, which is necessary in a typically timeboxed scenario of CTFs, trainings, or exams. The selected functionality also allows for efficiently storing C T F flags i n the note content and using note identifiers as flag IDs. Most importantly, the authorized data retrieval functionality: sharing and viewing content of personal or shared notes allows the presence of authorization flaws, and the overall feature design allows for a possible S Q L injection i n each functionality, as all of it is designed to be database-centric. The template is built to be flexible and easy to modify and extend, m a k i n g it simple to adapt to different database systems, frontend styles, and vulnerability types. We have been following Squarcina's A / D Service Development Guidelines [34] to make the generated application enjoyable to exploit 1. In this setting, transpilation denotes the automatic translation of a query from one SQL dialect to another (e.g., PostgreSQL -> MySQL) while preserving its semantics. 22 3. D E S I G N and patch. The guidelines also helped us design the generated application to be resilient against different variations of Denial of Service attacks. These attacks happen both intentionally and unintentionally during an A / D CTF, and they are often not preventable by the rules of the competition. The typical unintentional scenario is a h i g h load caused by a large number of teams exploiting a single service. O r an inefficient exploit that consumes a significant part of the resources available to the C T F service. The generated application can also be used outside of the A / D environment, for example, as a standalone challenge in a typical jeopardystyle CTF. Configuration generator Since the configuration of the generator core requires the specification of numerous interrelated parameters, we introduce a layer responsible for creating the configuration file for the generator core from user input. It works by accessing a configuration file w i t h vulnerability definitions for a given Application Baseline. The configuration generator can also list the available vulnerabilities, descriptions, and unique identifiers. The configuration generator takes the following input: • Vulnerability specification: - Number, type, and difficulty of the vulnerabilities - Unique identifier of a vulnerability • D B M S specification (optional) • Stylesheet specification (optional) After processing the user input, the generator determines a set of vulnerabilities to be used. It selects the optimal configuration and D B M S based on the requirements of all vulnerabilities, as some v u l nerabilities may require specific conditions i n order to be exploitable. The output of this module is a configuration file that can be directly supplied to the generator core. A n overview of this architecture is presented i n Figure 3.1. 23 4 Implementation This chapter describes h o w the vulnerable service generator is implemented. It first covers the core functionality, illustrating how vulnerabilities are dynamically inserted into an existing application template, and explains the technical details behind the vulnerabilities and associated hardening mechanisms. Secondly, it introduces a configuration generator that translates user input into ready-to-use configuration files. Next, it describes the baseline code structure a n d technology stack of the generated vulnerable application. The chapter then discusses the validation scripts produced alongside the application. Finally, it explains h o w the entire generation-to-deployment process is integrated into an automated workflow. The chapter concludes w i t h an estimation of how many distinct applications the system can generate, considering combinations of vulnerabilities, hardening strategies, and randomization of visual and structural elements. 4.1 Generator core The core functionality of the generator is implemented i n Python, w h i c h leverages the flexibility of Jinja2 for template processing a n d code generation. The implementation employs an object-oriented approach through the central VulnerableQueryGenerator class that orchestrates the entire generation process. Configuration and vulnerability processing The generator begins by loading the vulnerability configuration from a Y A M L file created by the configuration generator. This configuration specifies the database type, table names, a n d most importantly, the intended vulnerability locations. The generator extracts details about w h i c h components should contain vulnerabilities. SQL transpilation One of the features of the generator is its ability to support multiple database management systems. The implementation utilizes the 24 4. I M P L E M E N T A T I O N // Secure version using a prepared statement await query("SELECT * FROM users WHERE username = ? " , [username]); // Transformed vulnerable version using // string concatenation await query("SELECT * FROM users WHERE username = " ' + username + " ' " ) ; Listing 1: Comparison of secure (parameterized) and insecure (concatenated) S Q L queries SQLGlot library (Section 2.4) to transpile schemas between different S Q L dialects, ensuring that the generated database schema is compatible w i t h the selected database system. The dialects supported are SQLite, PostgreSQL, and M y S Q L . The original schemas and queries are written i n M y S Q L and automatically transpiled into the other dialects w h e n necessary. Query transformation The implementation transforms secure queries into vulnerable ones only at specified locations, applying these changes to the designated parameters w h i l e preserving the functionality of the application. A secure prepared statement [33] can be transformed into a vulnerable version using string concatenation, as illustrated i n Listing 1. SQL Injection Hardening Mechanisms The generator implements a hardening mechanism that modifies the difficulty of S Q L injection vulnerabilities by a d d i n g input filtering code as an ExpressJS middleware processing all relevant user input from an H T T P request. This middleware detects and blocks malicious input, but it also contains intentional weaknesses that can be bypassed through specific techniques listed below. 1. Backslash escape filter bypass [hard bypass difficulty] This filter blocks single quotes in user input. The filter can be bypassed 25 4. I M P L E M E N T A T I O N using backslash characters to escape quote delimiters. This technique works specifically i n scenarios w i t h multiple sequential parameters, such as: SELECT * FROM users WHERE username='userinputl' AND password='userinput2 1 A n attacker can inject m a l i c i o u s \ as u s e r i n p u t l . This causes the ' AND password=' portion to be interpreted as part of the first string literal. The userinput2 is then executed as S Q L syntax rather than treated as a string parameter, for example: SELECT * FROM users WHERE username='malicious\ 1 AND password=' OR 1=1—1 The database treats everything between the first quote and the quote after password= as a single string. This allows arbitrary S Q L injection i n the second parameter. However, this behavior is not consistent across the supported D B M S . According to the SQLite documentation, "C-style escapes using the backslash character are not supported because they are not standard S Q L " [35]. Therefore, SQLite does not support backslash escaping i n string literals. PostgreSQL supports backslash escapes only w h e n using a n explicit escape string syntax (E'...') [36]. In contrast, M y S Q L supports backslash escaping by default, as noted i n the M y S Q L reference manual [37]. Given that PostgreSQL requires a non-default syntax to enable this behavior, we considered this a n unrealistic scenario within the context of the vulnerable web application and its expected functionality. A s a result, we chose to support only the M y S Q L D B M S for this bypass mechanism. This resulted in limited D B M S variability o n this hardening mechanism i n order to provide a more realistic attack scenario. 2. Recursive filter bypass [medium bypass difficulty] This filter checks for S Q L keywords but only examines one level of input recursion. Attackers can bypass it b y w r a p p i n g parameters i n 26 4. I M P L E M E N T A T I O N s q l i : filter_enabled: true filter.patterns: ["SELECT", "UNION"," — " ] array_bypass: true uppercase_bypass: true Listing 2: Filter Y A M L configuration example arrays. The server processes these array parameters while the detection mechanism fails to identify the payload inside the array structure. Specifically, JavaScript's implicit type coercion converts array parameters to strings when concatenated, resulting i n the payload bypassing keyword checks. For instance, the expression ' c c ' + [' aa',' bb' ] evaluates to ' ccaa, bb', as arrays are implicitly joined into comma-separated strings [38]. Attackers leverage this behavior to evade detection mechanisms designed to inspect string inputs directly, thus bypassing the filter. 3. Uppercase filter bypass [medium bypass difficulty] This filter blocks S Q L keywords i n uppercase form only. Attackers can use lowercase or mixed-case variations such as select or SeLeCt instead of SELECT to bypass the filter. These methods can also be combined, resulting i n unique and more difficult-to-exploit scenarios. The filter configuration is controlled through the Y A M L configuration file, which, if needed, is automatically created by the configuration generator w h e n vulnerabilities of higher difficulty are requested. A configuration example is shown i n Listing 2. This configuration creates middleware that intercepts and filters requests before they reach vulnerable endpoints. SQL injection H a r d e n i n g selection The configuration generator conducts hardening selection during vulnerability specification processing. This selection determines appropriate S Q L injection protection mechanisms based o n the requested difficulty level and database compatibility requirements. A hardening 27 4. I M P L E M E N T A T I O N can also be used as a means to raise the difficulty level of a vulnerability with a lower difficulty. If enough vulnerabilities w i t h the requested difficulty are available, the generator may decide not to use any hardening mechanism. The system uses a compatibility-driven selection algorithm rather than a simple one-to-one vulnerability mapping. The algorithm follows these steps: 1. Filters vulnerabilities according to the requested type and difficulty level. 2. Identifies hardening mechanisms compatible with the requested difficulty 3. Creates groups of vulnerabilities based on hardening option compatibility. 4. Chooses an appropriate group using compatibility rules and randomization. 5. Verifies database system compatibility across all selected com- ponents. 6. Samples the required number of vulnerabilities from the compatible set. The output of this algorithm is a part of Y A M L configuration w i t h a set of vulnerabilities and the optionally selected hardening strategy as shown i n Listing 3. Authorization vulnerabilities The generator incorporates multiple authorization vulnerability types that exemplify common access control flaws. Each vulnerability represents a specific implementation error: 1. No return on failed response: This vulnerability occurs w h e n middleware sends an error response without terminating request processing. The code sends a 404 status code but lacks a r e t u r n statement before next() is called. Request execution continues despite the error condition, and protected resources become accessible. 28 4. I M P L E M E N T A T I O N s q l i : f i l t e r _ e n a b l e d : true f i l t e r _ p a t t e r n s : ["SELECT", " — ", "/*", "*/", "UNION"] array_bypass: true uppercase_bypass: f a l s e v u l n s : - component: g e t _ u s e r _ f o r _ s h a r i n g d i f f i c u l t y : medium i d : s q l i _ s h a r e _ u s e r name: error-based SQLi i n g e t _ u s e r _ f o r _ s h a r i n g parameters: - shareWithUsername t y p e : s q l i Listing 3: Example Y A M L output f r o m S Q L i hardening selection al- gorithm 2. Incorrect parameter order: This vulnerability occurs w h e n application code checks request parameters i n an improper sequence. The code examines r e q . body parameters before r e q . params. A t tackers can override U R L parameters by submitting H T T P body parameters alongside legitimate U R L parameters. Example code w i t h this behavior is shown i n Listing 4. If an adversary submits a request to the endpoint / p o s t / 1 , the framework correctly dispatches the call to the handler responsible for resource 1. Nevertheless, by embedding an H T T P b o d y content such as itemld=42, the handler selects the value 42 from the request b o d y i n preference to the U R L parameter. The application therefore retrieves a n d authorises access to post 42, disclosing data to which the requester is not entitled (post w i t h i d l ) . 3. Wrong parameter name checks: In this implementation, middleware examines incorrectly named parameters ( i t emId instead of i d or vice versa). The parameter name mismatch results i n null checks that always pass, w h i c h circumvents authorization checks for specific endpoints. 29 4. I M P L E M E N T A T I O N const id = req?.body?.itemld ?? req.params?.id; i f (!canAccess(id)){ return res.send(403) } Listing 4: Incorrect parameter precedence: body parameters overriding U R L parameters 4. Default allow on empty username: This flaw stems from i n complete authentication validation. The code verifies parameter existence but not user authentication state. W h e n a username is empty or n u l l , the middleware permits the request without proper authorization verification. These vulnerabilities are implemented as alternative middleware modules selected according to the configuration file. The generator replaces secure authentication middleware w i t h one of these flawed implementations. Each middleware file represents a distinct vulnerability variant. The middleware files are stored in the directory named /vuln_code/auth_middleware/. Authorization flaws differ from S Q L injections as they typically involve logic errors rather than database query manipulation. The provided Proof of Concept (PoC) scripts demonstrate methods to circumvent these controls and access restricted resources. Template rendering The final step i n the generation process is applying the Jinja2 [15] templates to create the complete application. The generator walks through all files i n the template directory, processing any file w i t h the . j 2 extension using Jinja2 templating. Files without this extension are copied directly to maintain directory structure and non-template assets. 30 4. I M P L E M E N T A T I O N 4.2 Configuration generator The configuration generator, implemented in generate_vuln_conf . py, converts user inputs into configuration files required by the generator core. C o m m a n d - l i n e interface The configuration generator accepts command-line arguments that specify the number, difficulty, and type of vulnerabilities to include or a specific identification of a requested vulnerability. Example usage: python generate_vuln_conf.py l : e a s y : s q l i 1:medium:auth python generate_vuln_conf.py v u l n : s q l i _ n o t e s _ p a r a m This approach removes the need for users to interact directly with the configuration file format of the generator core. Vulnerability definitions The configuration generator relies on the presence of a configuration file named v u l n s . yaml, which contains a structured list of predefined vulnerabilities, hardening methods, and stylesheets. This file was manually created and can be easily extended, allowing configurations compatible with the baseline application template to be added or modified as needed. The configuration generator uses this file to determine w h i c h vulnerabilities can be selected and supplied into the resulting configuration file for the generator core. Vulnerability selection logic This generator implements a selection algorithm for configuring v u l nerabilities based on user specifications. It filters available vulnerabilities by type and difficulty, and then groups them based on their compatibility w i t h various hardening mechanisms, ensuring compatibility among the selected vulnerabilities. 31 4. I M P L E M E N T A T I O N D B M S compatibility handling The configuration generator determines compatible database systems for the selected vulnerabilities by analyzing the D B M S requirements for each vulnerability and hardening mechanism, then computing the intersection of all compatibility sets. This process identifies database systems that can support all selected components. Configuration output The final output is a structured Y A M L configuration file containing specifications for the database system, table names, stylesheet, S Q L injection hardening mechanisms, and the selected vulnerabilities, as shown i n Listing 5. 4.3 Baseline code template The baseline code template, located i n the t e s t s e r v i c e / directory, functions as the foundation for all generated applications. This template implements a web application with features designed to demonstrate security vulnerabilities i n a contextual environment. Technology stack The baseline application is built using modern web technologies: • Backend: Node.js w i t h Expresses framework • Database: Compatible w i t h SQLite, PostgreSQL, and M y S Q L • Frontend: H T M L , CSS, and JavaScript w i t h templating support • Authentication: Session-based with optional middleware config- urations This technology stack was chosen for its widespread use i n web development and for being easy for the C T F players to understand. 32 4. I M P L E M E N T A T I O N config: database: mysql service_config: stylesheet_name: modern, ess tables: items_table: items shared_items_table: shared_items users_table: users s q l i : array_bypass: false filter_enabled: true filter_patterns: _ i i i i - /* - '*/' uppercase_bypass: false vulns: - component: check_shared_item_access d i f f i c u l t y : easy id: sqli_shared_access_id name: SQLi (auth bypass) via i d i n check_shared_item_access parameters: - i d type: sqli - component: auth_middleware d i f f i c u l t y : medium id: auth_bypass_incorrect_order method: incorrect_order name: Authentication bypass - incorrect order of using arguments type: auth Listing 5: Example Y A M L configuration file generated by the configuration generator 33 4. I M P L E M E N T A T I O N Itens Item Details Figure 4.1: User Interface Screenshots of the Application A p p l i c a t i o n structure The baseline application implements a note-sharing service w i t h the following features: • User registration and authentication • Creation and retrieval of personal notes • Sharing notes w i t h other users • Viewing shared notes from other users The application structure supports the systematic integration of security vulnerabilities within a coherent and realistic usage scenario. The user interface corresponding to this structure is s h o w n i n Figure 4.1. Database schema The baseline template defines a database schema that includes tables for users, items, and shared items. The schema is designed to be compatible w i t h multiple database systems while supporting the core functionality required for demonstrating SQL injection vulnerabilities. 34 4- I M P L E M E N T A T I O N // Template version (before transformation) const { username, password } = req.body; const row = await {{ query( "SELECT * FROM users WHERE username = ? AND password = ?", placeholder_vars=["username", "password"], query_id="login_check_user") }> // Transformed version (vulnerable to SQLi) const { username, password } = req.body; const row = await query( "SELECT * FROM users WHERE username = "' + username + "' AND password = ?", [password] ); Listing 6: Template marker and its transformed (vulnerable) version Template markers Throughout the baseline code, special template markers w i t h function calls indicate where the generator should inject code or m o d i f y existing code. These markers use Jinja2 syntax to invoke the query transformation function from the generator core, allowing for precise control over where vulnerabilities are introduced. In the example s h o w n i n Listing 6, the template marker is the {{ query (. . .) }} block, w h i c h uses Jinja2 templating syntax {{}} to specify the structure of the query, the JavaScript variables to use i n the query, and the q u e r y _ i d which uniquely identifies the query enabling us to mark it as vulnerable as. In a configuration where the username parameter is vulnerable to S Q L i , the resulting code w o u l d look like shown i n Listing 6. API endpoints The baseline application defines RESTful A P I [39] endpoints that i m plement core functionality, including user registration, authentication, 35 4. I M P L E M E N T A T I O N and note operations. The generator selects endpoint U R L s randomly from a predefined wordlist rather than using fixed paths. This randomization creates variability between generated applications. Each instance receives unique A P I paths that differ from those of other generated instances. The wordlist can be extended by administrators who wish to add new path name options for each functionality type. A l l application components, including verification scripts and proof of concept exploits, receive these randomized endpoints and must adapt their requests accordingly. This prevents students w i t h prior knowledge of other generated services from relying on hardcoded paths during exploitation attempts. Stylesheets The baseline contains three different CSS stylesheets to visually distinguish the generated applications. A l l three stylesheets are being applied to the same underlying H T M L code structure, but create different visual appearances. This design allows the generated applications to have distinct looks without requiring changes to the functional H T M L structure. The available themes include: • 90s: Retro design w i t h bright colors and animated effects • M o d e r n : Clean m i n i m a l interface w i t h subtle shadows • Cyber: Dark theme w i t h neon text and terminal aesthetics The stylesheet selection is determined during the generation process, either randomly if not specified or based on user preference through a command line argument. This visual differentiation serves several purposes in educational settings: it helps distinguish between different exercises, reduces student recognition of previously seen applications, and increases the perceived variety of generated services. 4.4 Validation scripts The validation scripts, located i n t e s t _ s e r v i c e / under checker and poc directories, evaluate both functionality and exploitability of gener- 36 4. I M P L E M E N T A T I O N ated applications. These scripts verify that generated services operate correctly within a Capture The Flag environment. Checker scripts The checker/basic_checker. py script tests the functional correctness of generated applications. It performs tests of application features from user registration to note sharing. This script is used in the Gitlab C I / C D pipeline to verify the functionality of the application. The checkerbot script in checker/checker. py implements the FAUST A / D c h e c k e r l i b interface [13]. This interface allows the placement and verification of flags w i t h i n the application and dynamically tests the functionality. We have decided to duplicate the testing functionality of the FAUST checkerbot into the basic_checker script, as the c h e c k e r l i b library, on w h i c h the checkerbot relies, is not available i n the testing environment. The checkerbot also requires a connection to a running FAUST server to test the service's functionality successfully. M o c k i n g this library w o u l d be possible, but it w o u l d require significantly more effort to implement compared to a rewrite into a simple Python script. The basic_checker script is designed to be r u n locally, without the need to connect to the FAUST server. Depending on the user's preference, it can be r u n i n a Docker container or directly on the host machine. This also makes implementing manual changes in the application code easier because it allows testing the generated application locally before deploying it into a production environment. Proof of Concept (PoC) scripts The templates of PoC scripts in t e s t _ s e r v i c e / p o c / show exploitation methods for the generated vulnerabilities. W h e n a new vulnerable application is generated, these templates are used to generate P o C scripts that w i l l be able to demonstrate and therefore verify the i n tended vulnerabilities in the generated application. These scripts serve two purposes - they validate vulnerability presence and function as educational material for the students w h o d i d not manage to find the solution during the exercise. 37 4. I M P L E M E N T A T I O N SQL Injection The S Q L injection P o C script ( p o c / s q l i .py) creates H T T P request files w i t h pre-marked injection points according to the vulnerability configuration. These request files are then used w i t h the S Q L M a p tool to verify the exploitability of the vulnerabilities. D e p e n d i n g on the configuration, this script also incorporates the use of S Q L M a p tamper scripts to bypass an input filter and the use of the array bypass technique mentioned i n Section 4.1. The output of this script is a m a r k d o w n file containing the H T T P request along w i t h the payload used by S Q L M a p i n order to exploit the injection. A t the current state, this script can verify the exploitability of firstorder injections i n H T T P P O S T parameters and H T T P path parts. It can also apply measures to bypass the input filtering for the Recursive and Uppercase filter bypasses described i n Section 4.1. The support for second-order S Q L i and bypassing the Backslash escape filter is planned as future work i n Chapter 6. Authorization flaws Unlike the S Q L i , the authorization flaw PoCs are mostly static. They usually require only minor changes after generation, such as updating H T T P path names. PoC for each authorization vulnerability contains a set of helper functions that manage making H T T P requests to the application. The PoCs then use these functions to exploit the vulnerability. A snippet of such PoC is shown i n Listing 7. 4.5 Automation The infrastructure, centered around the GitLab C I / C D [22] pipeline, automates the process of generating, building, deploying, and testing the vulnerable applications. C I / C D pipeline structure The GitLab CI pipeline consists of a main generate job that performs four logical phases operating in sequence. This design choice of having four logical phases within a single CI stage rather than separate stages 38 4. I M P L E M E N T A T I O N # Create sessions for each user sessionl = requests . SessionO sessions = requests . SessionO session3 = requests . SessionO # Register and login first user # Share the item with the second user share_item(sessionl, item_id, user2) # Register and login third user user3 = generate_random_string() password3 = generate_random_string() register(session3, user3, password3) login(session3, user3, password3) # Use the "no_return" vulnerability to # share the item with the third user share_item(session3, item_id, user3) # Verify that the third user has the shared item shared_items = get_shared_items(session3) assert any(item["id"] == item_id for item i n shared_items), "Shared item not found for third user" Listing 7: Snippet from no_return. py. j 2 with example usage of the helper functions 39 4. I M P L E M E N T A T I O N stems from several important considerations: the tight sequential dependency between phases where each phase requires immediate outputs from the previous one, shared environment requirements across all phases, the need for atomic success or failure of the entire process, management and communication w i t h ephemeral Docker resources that exist only during job execution, and simplified pipeline logic for comprehensive error handling and maintenance. The pipeline also includes a documentation job (push_available_vulns) that automatically executes whenever code is pushed to the repository, generating and updating a list of all available vulnerability configurations ( a v a i l a b l e _ v u l n s .md for users to reference. Configuration Generation Phase This initial phase processes userspecified vulnerability parameters supplied through the VULN_SPECS environment variable. The content of this variable is directly used as the input for the configuration generator to create a structured Y A M L specification that guides the subsequent generation process. The configuration defines w h i c h vulnerabilities to include, their difficulty levels, and any special requirements such as database compatibility. Service Generation Phase In this phase, the pipeline takes the configuration file and transforms it into a complete application codebase. The generator core parses the template files and injects vulnerabilities at predefined locations. This process includes transpiling database schemas, randomizing endpoint names, and implementing the specified vulnerabilities according to the configuration created i n the previous step. Validation Phase The validation phase builds the application within Docker containers and subjects it to automated testing. The pipeline executes both basic functionality tests basic_checker. py and targeted vulnerability checks located i n generated_service/poc, failing if any critical issues are detected. A s shown i n Figure 4.2, this phase follows the service generation phase and transitions directly into building the Docker containers w i t h i n the C I / C D pipeline. 40 4. I M P L E M E N T A T I O N CYBERSEC / theses / geleta-v Search visible log output | Q j [ g| ] rp | t Downloading argparse-1.4.3-py2.py3-none-any.whl (23 kB) Downloading requests-2.32.3-py3-none-any.ivhl (.64 kB) Downloading certifi-2025.1.31-py3-none-any.vihl (166 kB) Downloading charset_narmalizer-3.4.1-cp312-cn312-musllinux_l_2_x86_64.whl (146 kB) Downloading idna-3.ia-py3-none-any.whl (79 kB) Downloading MarkupSafe-3.a.2-cp312-cp312-musllinux_l_2_>:86_fiii.whl [23 kB) Downloading urllib3-2.3.9-py3-none-any.whl (128 kB) Installing collected packages: argparse, urllib3, sqlglot, pyyaml, HarkupSafe, idna, certifi, requests, ]inja2 Successfully installed MarkupSafe-3.3.2 argparse-l.