首页期刊视频编委会征稿启事出版道德声明审稿流程读者订阅论文查重联系我们English
引用本文
  • 李家强,李淑君,喻庞泽,等.基于改进PPO的雷达抗干扰自适应跳频方法[J].电讯技术,2026,66(9): - .    [点击复制]
  • LI Jiaqiang,LI Shujun,YU Pangze,et al.A Radar Anti-jamming Adaptive Frequency Hopping Method Based on Improved Proximal Policy Optimization(PPO) Algorithm[J].,2026,66(9): - .   [点击复制]
【HTML】 【打印本页】 【下载PDF全文】 查看/发表评论下载PDF阅读器关闭

←前一篇|后一篇→

过刊浏览    高级检索

本文已被:浏览 28次   下载 6 本文二维码信息
码上扫一扫!
基于改进PPO的雷达抗干扰自适应跳频方法
李家强,李淑君,喻庞泽,陈金立,姚昌华
(南京信息工程大学 电子与信息工程学院,南京 210044)
摘要:
针对现代电子战中雷达系统面临的日益复杂干扰问题,提出了一种基于深度强化学习的频率捷变雷达智能抗干扰决策算法。通过将深度强化学习方法与频率捷变雷达相结合,提出了一种基于近端策略优化的雷达自适应跳频抗干扰方法。首先,将雷达和干扰机之间的相互作用过程建模为部分可观测马尔可夫决策过程,基于检测概率设计了奖励函数。然后,采用近端策略优化(Proximal Policy Optimization,PPO)算法,通过其裁剪操作来解决传统策略梯度算法中策略更新过大的问题。最后,引入多层历史观测信息,并使用门控循环单元来提取序列数据中的长期依赖关系,以解决部分马尔可夫决策过程中信息不足的问题。仿真结果表明,该方法相较于传统近端策略优化算法能够将检测概率提高8%。
关键词:  频率捷变雷达  雷达抗干扰  自适应跳频  深度强化学习  近端策略优化(PPO)
DOI:10.20079/j.issn.1001-893x.250415006
基金项目:
A Radar Anti-jamming Adaptive Frequency Hopping Method Based on Improved Proximal Policy Optimization(PPO) Algorithm
LI Jiaqiang,LI Shujun,YU Pangze,CHEN Jinli,YAO Changhua
(School of Electronics and Information Engineering,Nanjing University of Information Science and Technology,Nanjing 210044,China)
Abstract:
In modern electronic warfare,radar systems face increasingly complex jamming challenges.An intelligent anti-jamming decision-making algorithm for frequency-agile radar based on deep reinforcement learning(DRL) is proposed.By integrating DRL with frequency-agile radar technology,an adaptive frequency hopping anti-jamming method is proposed using the proximal policy optimization(PPO) algorithm.First,the interaction between the radar and the jammer is modeled as a partially observable Markov decision process(POMDP),and the reward function is designed based on detection probability.Then,the PPO algorithm is employed,utilizing its clipping mechanism to address the issue of excessive policy updates in traditional policy gradient methods.Finally,multi-layer historical observations are incorporated,and a gated recurrent unit(GRU) is applied to capture long-term dependencies in sequential data,mitigating the information insufficiency in POMDP.Simulation results show that,compared with traditional PPO algorithms,the proposed method improves the detection probability by 8%.
Key words:  frequency-agile radar  radar anti-jamming  adaptive frequency hopping  deep reinforcement learning  proximal policy optimization(PPO)
安全联盟站长平台