یادگیری تقویتی چارچوبی به‌منظور پهنه‌بندی سیل در مناطق کوهستانی با استفاده از مدل Deep Q-Learning

نوع مقاله : پژوهشی

نویسندگان

1 استادیار بخش تحقیقات حفاظت خاک و آبخیزداری، مرکز تحقیقات و آموزش کشاورزی و منابع طبیعی استان چهارمحال و بختیاری، سازمان تحقیقات، آموزش و ترویج کشاورزی، شهرکرد، ایران

2 محقق بخش تحقیقات حفاظت خاک و آبخیزداری، مرکز تحقیقات و آموزش کشاورزی و منابع طبیعی استان چهارمحال و بختیاری، سازمان تحقیقات، آموزش و ترویج کشاورزی، شهرکرد، ایران

3 مربی بخش تحقیقات حفاظت خاک و آبخیزداری، مرکز تحقیقات و آموزش کشاورزی و منابع طبیعی استان چهارمحال و بختیاری، سازمان تحقیقات، آموزش و ترویج کشاورزی، شهرکرد، ایران

10.22092/wmrj.2026.371991.1656

چکیده

مقدمه و هدف
درک خطر سیلاب‌ها و تلاش برای کاهش آن از اولویت‌های اصلی پژوهشگران و سیاست‌گذاران است. پهنه‌بندی حساسیت به سیلاب ابزار مهمی برای مدیریت خطر در مناطق کوهستانی به‌شمار می‌آید. زیرا، می‌توان مناطق آسیب‌پذیر را شناسایی کرد و بستر برنامه‌ریزی مکانی، آمادگی و اقدامات پیشگیرانه را فراهم آورد. رویکردهای سنتی نقشه‌برداری سیلاب مانند مدل‌سازی آب‌شناختی و تحلیل داده‌های تاریخی اغلب با محدودیت‌هایی چون کمبود داده، پیچیدگی محاسباتی و پویایی ساختار‌های محیطی روبه‌رو هستند. ازاین‌رو، روش‌های نوین هوش مصنوعی به‌ویژه یادگیری ماشین و یادگیری تقویتی برای پردازش داده‌های گوناگون و شناسایی الگوهای پیچیده به‌کار گرفته شده‌اند. در میان رویکردهای هوش مصنوعی، یادگیری تقویتی به‌دلیل قابلیت تعامل با محیط و بهبود تدریجی تصمیم‌گیری، نویدبخش است. با به‌کارگیری مدل Deep Q-Learning (DQL) به‌عنوان یکی از روش‌های پیشرفته یادگیری تقویتی، می‌توان امکان کار با فضاهای حالت پیوسته و ابعاد بزرگ بدون نیاز به گسسته‌سازی را فراهم آرود. این پژوهش با هدف بهره‌گیری از مدل DQL برای تهیه نقشه حساسیت به سیلاب در استان چهارمحال و بختیاری با مناطق کوهستانی نیمه‌خشک و آسیب‌پذیر از سیلاب‌های ناگهانی ناشی از ذوب برف و بارش‌های شدید، انجام ‌شد. هدف اصلی، ارزیابی خطر سیلاب و ارائه چارچوبی دقیق برای سیاست‌گذاران به‌منظور مدیریت بحران، کاهش خطر و بهبود تاب‌آوری به‌ویژه در شهرستان‌های آسیب‌پذیر همچون اردل و لردگان بود.
مواد و روش‌ها
برای انجام این پژوهش، فهرست رخدادهای سیل شامل ۵۴۵ نقطه مکانی (۳۴۶ نقطه با سیلاب و ۱۹۹ نقطه بدون سیلاب) در بازه زمانی ۱۹۸۳ تا ۲۰۲۳ گردآوری شد. به‌طور تصادفی و لایه‌ای، ۷۱% از داده‌ها برای آموزش (۳۸۹ نقطه) و 29% آن برای آزمون (۱۵۶ نقطه) استفاده شد. ابتدا ۲۱ متغیر محیطی از منابع مختلف شامل مدل رقومی ارتفاع، سری زمانی اقلیمی، تصویرهای لندست 9 و پایگاه SoilGrid استخراج شد. پس از انجام تحلیل هم‌خطی، چهار متغیر حذف و ۱۷ متغیر نهایی (بلندی، تجمع جریان، جهت جریان، پیوستگی جریان، جهت جغرافیایی، طول شیب، انحنای افقی، نیم‌رخ، شاخص رطوبت پستی‌بلندی، ژرفا تا بستر سنگی، بیشینه بارش ۲۴ ساعته، دماهای میانگین و بیشینه، ژرفای برف، NDVI، درصد ماسه سطحی و فاصله از مناطق مسکونی) انتخاب شد. از مدل Deep Q-Learning برای پردازش متغیرهای پیوسته و تولید نقشه حساسیت به سیلاب استفاده شد. این نقشه به پنج رده (خیلی‌کم تا خیلی‌زیاد) طبقه‌بندی شد. دقت مدل به‌وسیلة شاخص‌هایی مانند AUC، Kappa، Recall،Precision ، Specificity و LogLoss  ارزیابی شد.
نتایج و بحث
نقشه طبقه‌بندی با استفاده از مدل Deep Q-Learning (DQL)  و بر اساس روش Natural break تهیه شد. در این نقشه، سطح منطقه مطالعه‌شده (۱۶۵۵۳۰۰ هکتار) به پنج طبقة حساسیت به سیلاب شامل خیلی‌کم (%۶، برابر با ۹۷۲۲۱ هکتار)، کم (%۲۰، برابر با ۳۳۶۱۹۰ هکتار)، متوسط (%۲۶، برابر با ۴۲۹۴۵۳ هکتار)، زیاد (%۲۶، برابر با ۴۳۵۲۹۳ هکتار) و خیلی‌زیاد (%۲۲، برابر با ۳۵۷۱۴۰ هکتار) تقسیم شد. شاخص‌های ارزیابی دقت، نشان‌دهنده کارایی بسیارخوب مدل بودند. نتایج شاخصAUC  (0/93) بیانگر توانایی تفکیک عالی، شاخص Kappa (0/72) بیانگر توافق قابل‌قبول میان پیش‌بینی و واقعیت، شاخصRecall  (0/87) بیانگر شناسایی زیاد نقاط سیلابی، بود. شاخص Precision برابر با 900/90 ، Specificity برابر با 0/85 و LogLoss برابر 0/65، محاسبه شد. تحلیل اهمیت متغیرها نشان داد بیشترین اثر مربوط به ژرفای برف بود که بر ذوب برف در سیلاب‌های بهاره بسیار اثرگذار است. پس ‌از آن تجمع جریان (ویژگی‌های پستی‌بلندی)، فاصله از مناطق مسکونی (اثرات انسانی و توسعه شهری)، شاخصNDVI (پوشش گیاهی و نفوذپذیری)، دمای میانگین (شرایط اقلیمی) و طول شیب بودند. مناطق با حساسیت خیلی‌زیاد عمدتاً در نواحی پست، دره‌ها و امتداد رود‌های اصلی مانند کارون و زاینده‌رود به‌ویژه در شهرستان‌های اردل و لردگان بودند. مقایسه نتایج این پژوهش با پژوهش‌های مشابه مبتنی بر مدل‌های یادگیری ماشین نظارت‌شده نشان داد که با کاربرد مدل DQL می‌توان خطر سیلاب در محیط‌های ناهمگن کوهستانی را به‌طور قابل قبول‌‌تری پیش‌بینی کرد و خطر سیلاب را با دقت بیشتری مدیریت کرد.
نتیجه‌گیری و پیشنهادها
نتایج این پژوهش بیانگر آن بود که مدل Deep Q-Learning (DQL) به‌عنوان یک روش یادگیری تقویتی عمیق، کارایی زیادی در پهنه‌بندی حساسیت به سیلاب به‌ویژه در مناطق پیچیده کوهستانی مانند استان چهارمحال و بختیاری دارد. با بهره‌گیری از این مدل و با پردازش متغیرهای محیطی پیوسته و شناسایی دقیق روابط غیرخطی، نقشه‌های معتبر و کاربردی تولید شد و بر اساس آن مناطق پرخطر به‌طور دقیق شناسایی شد. نتایج این پژوهش مؤید آن است که عامل‌های اصلی سیلاب در منطقه مطالعه‌شده شامل ذوب برف، ویژگی‌های پستی‌بلندی، توسعه انسانی و کاهش پوشش گیاهی است. این یافته‌ها با واقعیت‌های میدانی منطقه نیز هم‌راستا بود. از نقشه‌های تولیدشده می‌توان به‌عنوان الگوی عملی برای برنامه‌ریزی کاربری زمین، بهبود زیرساخت‌های دفع آب باران، حفاظت از مناطق مسکونی در برابر سیلاب‌های ناگهانی و بازسازی بوم‌سازگان‌های گیاهی به‌منظور افزایش نفوذپذیری خاک استفاده کرد. به‌طور کلی، نتایج این پژوهش رویکرد نوین و چارچوب دقیقی برای مدیریت پایدار خطر سیلاب در مناطق مشابه ایران و جهان ارائه داد که سیاست‌گزاران می‌توانند در بهبود تاب‌آوری جوامع محلی از آن بهره ببرند. برای بهبود دقت مدل در پژوهش‌های آینده، پیشنهاد می‌شود فهرست گسترده‌تری از رخدادهای سیلاب با داده‌های فصلی و زمانی دقیق‌تر گردآوری شود و متغیرهای پویا مانند تغییرات روزانه دما و چرخه‌های ذوب برف نیز بررسی شوند. افزون بر این، پیشنهاد می‌شود اعتبارسنجی میدانی گسترده‌تری به‌ویژه در بلندی‌های زیاد و مناطق شهری انجام شود. همچنین، ترکیب DQL با دیگر روش‌های یادگیری تقویتی یا مدل‌های ترکیبی به‌منظور افزایش کارایی آن پیشنهاد می‌شود.

کلیدواژه‌ها

موضوعات


عنوان مقاله [English]

A Reinforcement Learning Framework for Flood Zoning in Mountainous Areas using the Deep Q-Learning Model

نویسندگان [English]

  • Saleh Yousefi 1
  • Sara Mardanian 2
  • Ahmad Reza Karimipour 3
1 Assistant Professor in Soil Conservation and Watershed Management Research Department, Chaharmahal and Bakhtiari Agricultural and Natural Resources Research and Education Center, AREEO, Shahrekord, Iran
2 Scholar in Soil Conservation and Watershed Management Research Department, Chaharmahal and Bakhtiari Agricultural and Natural Resources Research and Education Center, AREEO, Shahrekord, Iran
3 Instructor in Soil Conservation and Watershed Management Research Department, Chaharmahal and Bakhtiari Agricultural and Natural Resources Research and Education Center, AREEO, Shahrekord, Iran
چکیده [English]

Introduction and Goal 
Understanding flood risk efforts to reduce it are top priorities for researchers and policymakers. Flood susceptibility mapping is an important tool for risk management in mountainous regions. Because, it is possible to identify prone areas and provide a platform for spatial planning, preparedness, and preventive measures. Traditional flood mapping approaches, such as hydrological modeling and historical data analysis, often face limitations including data scarcity, computational complexity, and the dynamic of environmental systems. Therefore, modern artificial intelligence (AI) methods—particularly machine learning (ML) and reinforcement learning (RL)—have been employed to process diverse datasets and identify complex patterns. Among AI approaches, reinforcement learning is promising due to its capacity to interact with the environment and progressively improve decision-making. The Deep Q-Learning (DQL) model, as an advanced reinforcement learning technique, enables operation in continuous and high-dimensional state spaces without requiring discretization. This study aims to apply the DQL model to produce a flood susceptibility map in the Chaharmahal and Bakhtiari Province, a mountainous and semi-arid region prone to flash floods caused by snowmelt and intense rainfall. The main objective is to provide a precise framework for flood hazard assessment and to support policymakers and crisis managers in reducing risk and enhancing resilience, particularly in vulnerable counties such as Ardal and Lordegan.
Materials and Methods
To conduct this research, a list of flood events including 545 spatial points (346 flood and 199 non-flood locations) from 1983 to 2023 was compiled. In a random and stratified manner, 71% of the data was used for training (389 points) and 29% for testing (156 points). Initially, 21 environmental variables were extracted from multiple sources, including a Digital Elevation Model (DEM), climatic time series, Landsat 9 imagery, and the SoilGrid database. After multicollinearity analysis, four variables were removed, and 17 final variables (elevation, flow accumulation, flow direction, stream connectivity, aspect, slope length, plan and profile curvature, Topographic Wetness Index (TWI), depth to bedrock, maximum 24-hour precipitation, mean and maximum temperature, snow depth, NDVI, surface sand percentage, and distance from residential areas) were selected. The Deep Q-Learning model was applied to process continuous variables and generate a flood susceptibility map. This map was classified into five categories (very low to very high). The accuracy of the model was evaluated by indicators such as AUC, Kappa, Recall, Precision, Specificity, and LogLoss metrics.
Results and Discussion
The classification map was prepared using the Deep Q-Learning (DQL) model and based on the Natural break method. In this map, the area of the study area (1,655,300 hectares) was divided into five flood susceptibility classes: very low (6%, 97,221 ha), low (20%, 336,190 ha), moderate (26%, 429,453 ha), high (26%, 435,293 ha), and very high (22%, 357,140 ha). Accuracy assessment indicated very good model performance. The results of the AUC index (0.93) indicated excellent discriminative ability, the Kappa coefficient index (0.72) indicating acceptable agreement between predictions and observations, and the Recall index (0.87) indicated high identification of flood points. The Precision index and was calculated to be 0.90, Specificity was 0.85, and LogLoss was 0.65. Analysis of the significance of the variables showed that the greatest effect was related to snow depth, which has a significant impact on snowmelt in spring floods. This was followed by flow accumulation (topographic characteristics), distance from residential areas (human and urban development impacts), NDVI (vegetation cover and soil permeability), mean temperature (climatic conditions), and slope length. Areas classified as very high susceptibility were mainly concentrated in low-lying areas, valleys, and along major rivers such as the Karoon and Zayandeh-Rud Rivers, particularly within Ardal and Lordegan counties. Comparison the results of this study with similar studies based on supervised machine learning models showed that using the DQL model, flood risk in heterogeneous mountainous environments can be predicted more reliably and flood risk can be managed more accurately.
Conclusion and Suggestions
The results of this study demonstrated that the Deep Q-Learning (DQL) model, as a deep reinforcement learning approach, has high capability for flood susceptibility mapping, especially in complex mountainous regions such as Chaharmahal and Bakhtiari Province. By utilizing this model, processing continuous environmental variables and accurately identifying nonlinear relationships, valid and practical maps were produced, and based on that, high-risk areas were accurately identified. The findings of this study confirm that the main flood drivers in the region include snowmelt, topographic characteristics, human development, and vegetation degradation. These findings were also consistent with the field realities of the region. The generated maps can serve as a practical basis for land-use planning, strengthening stormwater infrastructure, protecting residential areas against flash floods, and restoring vegetation ecosystems to enhance soil permeability. Overall, the results of this study provided a novel approach and a detailed framework for sustainable flood risk management in similar regions in Iran and the world, which policymakers can benefit from in improving the resilience of local communities. To improve the accuracy of the model in future research, it is recommended that a broader list of flood events be compiled with more accurate seasonal and temporal data, and that dynamic variables such as daily temperature changes and snowmelt cycles be examined. Additionally, it is recommended to conduct more extensive field validation, especially in high altitudes and urban areas. It is also suggested to combine DQL with other reinforcement learning methods or hybrid models to increase its efficiency.

کلیدواژه‌ها [English]

  • Flood risk assessment
  • Iranian highlands
  • machine learning
  • reinforcement learning
Alsumayt A, El-Haggar N, Amouri L, Alfawaer ZM, Aljameel SS. 2023. Smart flood detection with AI and Blockchain integration in Saudi Arabia using drones. Sensors. 23: 5148. https://doi.org/10.3390/s23115148
Arabameri A, Seyed Danesh A, Santosh M, Cerda A, Chandra Pal S, Ghorbanzadeh O, Roy P, Chowdhuri I. 2022. Flood susceptibility mapping using meta-heuristic algorithms. Geomatics, Natural Hazards and Risk. 13: 949–974. https://doi.org/10.1080/19475705.2022.2060138
Ardali EO, Tahmasebi P, Bonte D, Milotić T, Pordanjani IR, Hoffmann M. 2015. Ecological sustainability in rangelands: The contribution of dung beetles in secondary seed dispersal (Case study: Chaharmahal and Bakhtiari Province, Iran). European Journal of Sustainable Development. 5: 133. https://doi.org/10.14207/ejsd.2016.v5n3p133
Batzer DP, Sharitz RR. 2007. Ecology of freshwater and estuarine wetlands, Ecology of Freshwater and Estuarine Wetlands. Univ of California Press. 375 p. https://doi.org/10.1525/california/9780520247772.003.0001
Bowes BD, Tavakoli A, Wang C, Heydarian A, Behl M, Beling PA, Goodall JL. 2021. Flood mitigation in coastal urban catchments using real-time stormwater infrastructure control and reinforcement learning. Journal of Hydroinformaticss. 23: 529–547. https://doi.org/10.2166/HYDRO.2020.080
Bray J, Wasson RJ, Srivastava P, Ziegler AD. 2023. Floods and debris flows in Ladakh: Past history and future hazards, in: Advances in Asian Human-Environmental Research. Springer. pp. 31–52. https://doi.org/10.1007/978-3-031-42494-6_3
Chen W, Li Y, Xue W, Shahabi H, Li S, Hong H, Wang X, Bian H, Zhang S, Pradhan B, Ahmad B, Bin H. 2020. Modeling flood susceptibility using data-driven approaches of naïve Bayes tree, alternating decision tree, and random forest methods. Sci. Total Environ. 701: 134979. https://doi.org/10.1016/j.scitotenv.2019.134979
da Silva AR. 2022. Double Q-Learning for citizen relocation during natural hazards. arXiv Prepr. arXiv2209.03800.
Dewan AM. 2013. Floods in a megacity: Geospatial techniques in assessing hazards, risk and vulnerability, Floods in a Megacity: Geospatial Techniques in Assessing Hazards, Risk and Vulnerability. Springer. https://doi.org/10.1007/978-94-007-5875-9
Feng K, Lin N, Kopp RE, Xian S, Oppenheimer M. 2025. Reinforcement learning–based adaptive strategies for climate change adaptation: An application for coastal flood risk management. Proceedings of the National Academy of Sciences, 122(12), e2402826122.‏
Foy, P., 2019. Deep Reinforcement Learning: Guide to Deep Q-Learning, in: Mlq. Springer. pp. 475–483
Jaafari A, Mafi-Gholami D, Yousefi S. 2024. A spatiotemporal analysis using expert-weighted indicators for assessing social resilience to natural hazards. Sustainable Cities and Society. 100: 105051
Hall J, Arheimer B, Borga M, Brazdil R, Claps P, Kiss A, Kjeldsen TR, Kriauciuniene J, Kundzewicz ZW, Lang M, Llasat MC, Macdonald N, McIntyre N, Mediero L, Merz B, Merz R, Molnar P, Montanari A, Neuhold C, Parajka J, Perdigao RAP, Plavcova L, Rogger M, Salinas JL, Sauquet E, Schar C, Szolgay J, Viglione A, Bloschl G. 2014. Understanding flood regime changes in Europe: A state-of-the-art assessment. Hydrology and Earth System Sciences. 18: 2735–2772. https://doi.org/10.5194/hess-18-2735-2014
Khan TA, Shahid Z, Alam M, Su’ud MM, Kadir K. 2019. Early flood risk assessment using Machine Learning: A Comparative study of SVM, Q-SVM, K-NN and LDA, in: MACS 2019 - 13th International Conference on Mathematics, Actuarial Science, Computer Science and Statistics, Proceedings. IEEE, pp. 1–7. https://doi.org/10.1109/MACS48846.2019.9024796
Kharazmi R, Tavili A, Rahdari MR, Chaban L, Panidi E, Rodrigo-Comino J. 2018. Monitoring and assessment of seasonal land cover changes using remote sensing: A 30-year (1987–2016) case study of Hamoun Wetland, Iran. Environmental Monitoring and Assessment. 190: 1–23. https://doi.org/10.1007/s10661-018-6726-z
Kondolf GM, Piégay H, Landon N. 2007. Changes in the riparian zone of the lower Eygues River, France, since 1830. Landscape Ecology. 22: 367–384. https://doi.org/10.1007/s10980-006-9033-y
Li XY, Wang X. 2025. Rescue path planning for urban flood: A deep reinforcement learning–based approach. Risk Analysis. 45(4): 928-943.‏
Mohammadi M, Abdi Y, Rahimian FP. 2025. Blockchain for water management, in: Digital Twin and Blockchain for Sensor Networks in Smart Cities. Elsevier. pp. 229–242. https://doi.org/10.1016/B978-0-443-30076-9.00011-X
Naghibi SA, Pourghasemi HR. 2015. A comparative assessment between three machine learning models and their performance comparison by bivariate and multivariate statistical methods in groundwater potential mapping. Water Resources Management. 29: 5217–5236. https://doi.org/10.1007/s11269-015-1114-8
Pal S, Talukdar S. 2018. Application of frequency ratio and logistic regression models for assessing physical wetland vulnerability in Punarbhaba river basin of Indo-Bangladesh. Human and Ecological Risk Assessment: An International Journal. 24: 1291–1311. https://doi.org/10.1080/10807039.2017.1411781
Pourghasemi HR, Gayen A, Edalat M, Zarafshar M, Tiefenbacher JP. 2019. Is multi-hazard mapping effective in assessing natural hazards and integrated watershed management? Geoscience Frontiers. 11(4): 1203-1217. https://doi.org/10.1016/j.gsf.2019.10.008
Qin H, Liang Q, Chen H, De Silva V. 2024. A coupled human and natural systems (CHANS) framework integrated with reinforcement learning for urban flood mitigation. Journal of Hydrology. 643: 131918. https://doi.org/10.1016/j.jhydrol.2024.131918
Remesan R, Bray M, Shamim MA, Han D. 2009. Rainfall-runoff modelling using a wavelet-based hybrid SVM scheme, in: IAHS-AISH Publication. IAHS Press. pp. 41–50.
Schmid U, Rösch P, Krause M, Harz M, Popp J, Baumann K. 2009. Gaussian mixture discriminant analysis for the single-cell differentiation of bacteria using micro-Raman spectroscopy. Chemometrics and Intelligent Laboratory Systems. 96: 159–171. https://doi.org/10.1016/j.chemolab.2009.01.008
Tan H. 2021. Reinforcement learning with deep deterministic policy gradient, in: Proceedings - 2021 International Conference on Artificial Intelligence, Big Data and Algorithms, CAIBDA 2021. IEEE, pp. 82–85. https://doi.org/10.1109/CAIBDA53561.2021.00025
Tian W, Xin K, Zhang Z, Zhao M, Liao Z, Tao T. 2023. Flooding mitigation through safe and trustworthy reinforcement learning. Journal of Hydrology 620: 129435. https://doi.org/10.1016/j.jhydrol.2023.129435
Tian W, Zhang Z, Xin K, Liao Z, Yuan Z. 2025. Enhancing the resilience of urban drainage system using deep reinforcement learning. Water Research. 281: 123681.
Yousefi S, Jaafari A, Valjarević A, Gomez C, Keesstra S. 2022. Vulnerability assessment of road networks to landslide hazards in a dry-mountainous region. Environmental Earth Sciences. 81: 521. https://doi.org/10.1007/s12665-022-10650-z
Zhu X, Guo H, Huang JJ. 2024. Urban flood susceptibility mapping using remote sensing, social sensing and an ensemble machine learning model. Sustainable Cities and Society. 108: 105508. https://doi.org/10.1016/j.scs.2024.105508