Goodput maximization for large language model edge inference: A two-phase maskable PPO approach

Author Identifier (ORCID)

Wei Ni’s ORCID record ORCID Logo

Abstract

This letter presents a novel two-phase maskable proximal policy optimization (TP-MPPO) algorithm, which maximizes the system goodput counting request throughput with strict service level objective (SLO) compliance for large language model (LLM) inference services in wireless edge networks. In the first phase of TP-MPPO, we optimize the task offloading decisions by MPPO with action masking mechanism, effectively avoiding exploring invalid actions and reducing the action space. In the second phase, closed-form solutions are derived for uplink bandwidth allocation; a greedy algorithm is designed for downlink bandwidth allocation to provide immediate rewards for the MPPO in the next round. The two stages alternate till convergence. Simulation results demonstrate that TP-MPPO can improve the system reward by 33.3%-87.5% compared to its benchmarks and achieve the highest goodput.

Keywords

edge inference, large language model, resource allocation, task offloading

Document Type

Journal Article

Date of Publication

1-1-2026

E-ISSN

21622345

ISSN

21622337

Volume

15

Publication Title

IEEE Wireless Communications Letters

Publisher

IEEE

School

School of Engineering

Funding Information

Shanghai Municipal Science and Technology Commission Foundation (Grant Number: 25DP1500300 and 24DP1501001)

Copyright

subscription content

First Page

4400

Last Page

4404

Recommended Citation

Chen, X., Zhang, Q., Ni, W., Zhang, S., & Sun, Y. (2026). Goodput maximization for large language model edge inference: A two-phase maskable PPO approach. IEEE Wireless Communications Letters, 15, 4400–4404. https://doi.org/10.1109/LWC.2026.3718001

Share

 
COinS
 

Link to publisher version (DOI)

10.1109/LWC.2026.3718001