<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "https://jats.nlm.nih.gov/publishing/1.3/JATS-journalpublishing1-3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink"
         xmlns:mml="http://www.w3.org/1998/Math/MathML"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         article-type="research-article"
         dtd-version="1.3"
         xml:lang="en">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher-id">IJITEST</journal-id>
      <journal-title-group>
        <journal-title>International Journal of Innovative Trends in Engineering Science and Technology</journal-title>
        <abbrev-journal-title abbrev-type="publisher">IJITEST</abbrev-journal-title>
      </journal-title-group>
      <issn pub-type="epub">3139-6887</issn>
      <publisher>
        <publisher-name>Felix Academic Publications</publisher-name>
      </publisher>
      <self-uri xlink:href="https://ijitest.org"/>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="publisher-id">IJITEST-2026-013</article-id>
      <article-categories>
        <subj-group subj-group-type="heading">
          <subject>Original Research Articles</subject>
        </subj-group>
        <subj-group subj-group-type="article-type">
          <subject>Research Article</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Deep Reinforcement Learning for Dynamic Resource Management in Ephemeral Edge Computing Networks</article-title>
      </title-group>
      <contrib-group>
      <contrib contrib-type="author" corresp="yes">
        <name>
          <surname>Priya</surname>
          <given-names>Dr. Ch. Swapna</given-names>
        </name>
        <email>swapnachsp@gmail.com</email>
        <xref ref-type="aff" rid="aff1"/>
      </contrib>
      <contrib contrib-type="author">
        <name>
          <surname>Jani</surname>
          <given-names>Mahamed Mastan</given-names>
        </name>
        <email>mdjani1209@gmail.com</email>
        <xref ref-type="aff" rid="aff1"/>
      </contrib>
      <contrib contrib-type="author">
        <name>
          <surname>Mycherla</surname>
          <given-names>Bharath Karthik</given-names>
        </name>
        <email>bharathkarthik2006@gmail.com</email>
        <xref ref-type="aff" rid="aff1"/>
      </contrib>
      <contrib contrib-type="author">
        <name>
          <surname>Medisetty</surname>
          <given-names>Surya Teja</given-names>
        </name>
        <email>suryatejamedisetty000@gmail.com</email>
        <xref ref-type="aff" rid="aff1"/>
      </contrib>
      <contrib contrib-type="author">
        <name>
          <surname>Pulipati</surname>
          <given-names>Kalpana</given-names>
        </name>
        <email>pkalpana1109@gmail.com</email>
        <xref ref-type="aff" rid="aff1"/>
      </contrib>
      </contrib-group>
    <aff id="aff1">
      <institution-wrap>
        <institution content-type="orgname">Department of CSE, Vignan’s Institute of Information Technology (A), Visakhapatnam</institution>
      </institution-wrap>
    </aff>
      <pub-date date-type="pub" publication-format="electronic">
        <day>01</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <volume>1</volume>
      <issue>2</issue>
      <fpage>10</fpage>
      <lpage>15</lpage>
      <history>
        <date date-type="received" iso-8601-date="2026-05-02">
          <day>02</day>
          <month>05</month>
          <year>2026</year>
        </date>
        <date date-type="accepted" iso-8601-date="2026-08-01">
          <day>01</day>
          <month>08</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>Copyright &#169; 2026 Dr. Ch. Swapna Priya, Mahamed Mastan Jani, Bharath Karthik Mycherla, Surya Teja Medisetty, Kalpana Pulipati. Published by Felix Academic Publications.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <copyright-holder>Dr. Ch. Swapna Priya, Mahamed Mastan Jani, Bharath Karthik Mycherla, Surya Teja Medisetty, Kalpana Pulipati</copyright-holder>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are properly credited.</license-p>
        </license>
      </permissions>
      <self-uri content-type="html" xlink:href="https://ijitest.org/archives/volume1/issue2/IJITEST-2026-013"/>
      <self-uri content-type="pdf" xlink:href="https://ijitest.org/api/files/published/IJITEST-2026-013-published.pdf"/>
      <abstract xml:lang="en">
        <p>Efficient resource orchestration in modern edge computing deployments is increasingly challenged by node mobility, stochastic workloads, and limited energy budgets. Conventional static and heuristic scheduling methods are fundamentally inadequate for volatile environments such as UAV swarms and vehicular ad hoc networks, where topology and resource availability evolve continuously. This paper proposes a novel adaptive resource management framework grounded in Proximal Policy Optimization (PPO), a state-of-the-art Deep Reinforcement Learning (DRL) algorithm, tailored for ephemeral edge computing scenarios. The resource allocation problem is rigorously formalized as a Markov Decision Process (MDP) that jointly accounts for end-to-end task latency, cumulative energy expenditure, load distribution fairness, and Service Level Agreement (SLA) compliance. Through iterative interaction with a realistic simulation environment encompassing 20 mobile UAV nodes, the PPO agent acquires nuanced allocation policies that balance competing performance objectives. Our key novelty lies in a composite reward signal that explicitly penalizes battery depletion events, discouraging greedy local processing in favor of energy-balanced, network-lifetime-aware decisions. Experimental results demonstrate that the proposed PPO-based framework reduces SLA violations by approximately 30% and extends network operational lifetime by up to 47% compared to Deep Q-Network (DQN) baselines and classical static schedulers.</p>
      </abstract>
      <kwd-group kwd-group-type="author-keywords">
        <kwd>Deep Reinforcement Learning</kwd>
        <kwd>Proximal Policy Optimization (PPO)</kwd>
        <kwd>Edge Computing</kwd>
        <kwd>Dynamic Resource Allocation</kwd>
        <kwd>Markov Decision Process</kwd>
        <kwd>UAV Networks</kwd>
        <kwd>Energy Efficiency</kwd>
        <kwd>SLA Compliance</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-summary">
      <title>Article Overview</title>
      <p>Efficient resource orchestration in modern edge computing deployments is increasingly challenged by node mobility, stochastic workloads, and limited energy budgets. Conventional static and heuristic scheduling methods are fundamentally inadequate for volatile environments such as UAV swarms and vehicular ad hoc networks, where topology and resource availability evolve continuously. This paper proposes a novel adaptive resource management framework grounded in Proximal Policy Optimization (PPO), a state-of-the-art Deep Reinforcement Learning (DRL) algorithm, tailored for ephemeral edge computing scenarios. The resource allocation problem is rigorously formalized as a Markov Decision Process (MDP) that jointly accounts for end-to-end task latency, cumulative energy expenditure, load distribution fairness, and Service Level Agreement (SLA) compliance. Through iterative interaction with a realistic simulation environment encompassing 20 mobile UAV nodes, the PPO agent acquires nuanced allocation policies that balance competing performance objectives. Our key novelty lies in a composite reward signal that explicitly penalizes battery depletion events, discouraging greedy local processing in favor of energy-balanced, network-lifetime-aware decisions. Experimental results demonstrate that the proposed PPO-based framework reduces SLA violations by approximately 30% and extends network operational lifetime by up to 47% compared to Deep Q-Network (DQN) baselines and classical static schedulers.</p>
    </sec>
  </body>
  <back>
    <sec sec-type="declarations">
      <title>Declarations</title>
      <p>The authors declare that no competing interests exist in relation to this published work.</p>
    </sec>
  </back>
</article>