Secure Resource Allocation in IoT Environments Using Deep Reinforcement Learning-Based Adaptive Defenses
Keywords:
Internet of Things (IoT), Deep Reinforcement Learning (DRL), Dueling DQN, A3C, Adaptive Defense, Secure Resource Allocation, Edge Computing, Trust Management, Latency Optimization, Energy Efficiency, Cybersecurity, Policy Convergence, Multi-Objective RewardAbstract
The success of Internet of Things (IoT) networks has also created a major problem in resource provisioning, as these networks are becoming increasingly heterogeneous, dynamic, and distributed. Simple and heuristic-based allocation schemes have become obsolete in terms of fulfilling the collective goals of latency reduction, energy efficiency, and secure defense against malicious attacks. To overcome such limitations, this paper suggests a new Deep Reinforcement Learning (DRL) adaptive defense architecture designed to provide resources efficiently and securely in hostile IoT environments. The suggested system combines three major components, viz., Trust Monitoring Module (TMM), Resource Profiling Agent (RPA), and Network Security Orchestrator (NSO), each of which works within the edge and fog layers in cooperation with each other. These modules dynamically evaluate the trustworthiness of devices, keep track of resources, and coordinate real-time policy enforcement in situations of changing threats. Two enhanced DRL models, namely Dueling Deep Q-Network (Dueling DQN) and Asynchronous Advantage Actor-Critic (A3C), are tested in a simulated environment in which adversarial situations involving Distributed Denial of Service (DDoS), flooding, and spoofing attacks are replicated. Quantitative findings: the Dueling DQN comes out with better scores when it comes to important performance parameters: it delivers 89.7 percent packets, has an average latency of 92.3 ms, an energy efficiency of 362 mWh, and a convergence score of 0.88 on the trust score when compared to 83.5 percent packets, 108.2 ms average latency, 295 mWh energy efficiency, and 0.82 convergent score on the trust score of A3C. Dueling DQN further indicates 26.7 percent enhanced policy convergence, 22.7 percent greater energy efficiency, and 18.3 percent superior reward slope stability. These results speculate on the much higher versatility, trust affiliation, and appropriateness of the framework in accommodating the ensuing real-time, mission-centered endeavors of IoT implementation.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Soft Computing and Data Mining

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.









