Authors: Dr. Pankaj Malik, Tejasva Sharma, Tasneem Dewaswala, Viyaylaxmi Sharma, Akshat Jain, Vaibhav Sharma
Abstract: Liquidity stress management requires financial institutions to make sequential funding and asset-allocation decisions under uncertainty, where rare but severe shortfalls carry disproportionate cost. Conventional reinforcement learning (RL) agents optimize expected return and are therefore poorly suited to this setting, since they are indifferent to the shape of the return distribution and provide no explicit safety guarantee against tail losses. This paper proposes DSRL-IQN, a distributional safe reinforcement learning framework that combines an Implicit Quantile Network (IQN) with a Conditional Value-at-Risk (CVaR) risk estimator and an explicit safety layer for liquidity stress management. Rather than learning a single expected return, DSRL-IQN learns the full return distribution conditioned on state and action, from which a CVaR estimate of tail risk is derived at each decision step. A constraint-projection safety layer then restricts the agent to the subset of actions whose estimated CVaR remains within an institution-defined liquidity buffer, and selects the safe action with the highest expected return. The framework is evaluated on a liquidity-stress simulation environment constructed from balance-sheet and market-based liquidity indicators, under four synthetically generated stress regimes of increasing severity (Normal, Moderate, Severe, Extreme). DSRL-IQN is compared against DQN, Double DQN, PPO, QR-DQN and a plain IQN agent using return, Sharpe ratio, maximum drawdown, CVaR at the 5% level, and the rate of safety-constraint violations. Experimental results show that DSRL-IQN attains the highest risk-adjusted return and the lowest tail risk and violation rate across all stress regimes, while an ablation study confirms that the CVaR estimator and the safety layer contribute complementary and largely additive improvements. A sensitivity analysis further characterizes the trade-off between profitability and safety as the CVaR confidence level is varied. The results indicate that distributional, risk-constrained reinforcement learning is a practical and effective approach for liquidity stress management under uncertainty.
International Journal of Science, Engineering and Technology