Repository navigation
Retry exponential backoff max float overflow #47971
Description
Activity
- addedkind:bugThis is a clearly a bugThis is a clearly a bugneeds-triagelabel for new issues that we didn't triage yetlabel for new issues that we didn't triage yet
on Mar 19, 2025 It's a very, very, very niche case.
It's a very, very, very niche case.
Yes, but it is definitely a bug, which should be fixed.
I did new pull request: #48051Yes it is rare case, but It leads scheduler crash. From the provided configuration this failure will happen in ~41 days. If max_retry_delay will be 15 minutes it will be in 10+ days... Scheduler will be failed and you can't understand why without touching logs.
Yes it is rare case, but It leads scheduler crash. From the provided configuration this failure will happen in ~41 days. If max_retry_delay will be 15 minutes it will be in 10+ days... Scheduler will be failed and you can't understand why without touching logs.
Which does not make it realistic case to be honest. That's why it's super niche. You probably can get hundreds of unrealistic cases like that. And I am not saying it does not need to be fixed, it might, but looking at the PR, the code implemented there is far too long for the functionality.
Dear all,
Please see commit patch in /main, not only in 2.10.3
If it is not correct by form please do correct fix by yourself, if possible.Thank you in advance!
Dear @potiuk ,
What will be done with this bug? I provided different possible solutions for it, including simple patch to limit try_number to 500 in function, which calculates next retry delay.
It will be nice to fix this bug in upcoming releases.
I think it is better to catch the overflow and use the maximum delay from the task or the environment. However, I added the missing test cases to show the patch from alealandreev works and made a PR against their branch.
I submitted a PR against main that just catches the overflow and falls back to the maximum delay. No strong opinion about which is better.
Dear all,
Please see additional PR
PR 48557 is ready for review again.
I just had this issue in prod! is there a workaround to make it work? thanks god it just happened today in the morning! no one task was bing scheduled anymore! this needs to be fixed ASAP!
I addressed the comments on my PR which should solve the issue. You can back port it to whatever version you are using, @netogerbi.
Reacted by Reculos Gerbi Neto- added a commit that references this issue
on Aug 13, 2025 - added a commit that references this issue
on Aug 13, 2025 - added a commit that references this issue
on Aug 13, 2025 - added a commit that references this issue
on Aug 13, 2025 - added a commit that references this issue
on Sep 25, 2025 - added a commit that references this issue
on Oct 24, 2025 - added a commit that references this issue
on Mar 1, 2026 - added 2 commits that reference this issue
on Jul 30, 2026 - added a commit that references this issue
on Sep 8, 2026
Apache Airflow version
Other Airflow 2 version (please specify below)
If "Other Airflow 2 version" selected, which one?
2.10.3
What happened?
Hello,
I encountered with a bug. My DAG configs were: retries=1000, retry_delay=5 min (300 seconds), max_retry_delay=1h (3600 seconds). My DAG failed ~1000 times and after that Scheduler broke down. After that retries exceeded 1000 and stopped on 1017 retry attempt.
I did my research on this problem and found that this happened due to formula min_backoff = math.ceil(delay.total_seconds() * (2 ** (self.try_number - 1))) in taskinstance.py file. So retry_exponential_backoff has no limit of try_number and during calculations it can overflow max Float value. So even if max_retry_delay is set formula is still calculating. And during calculations on very large retry number it crashes.
Please fix bug.
I also did pull request with my possible solution:
#48057
#48051
From Airflow logs:
2024-12-09 02:16:39.825 OverflowError: cannot convert float infinity to integer
2024-12-09 02:16:39.825 min_backoff = int(math.ceil(delay.total_seconds() * (2 ** (self.try_number - 2))))
2024-12-08 09:29:14.583 [2024-12-08T06:29:14.583+0000] {scheduler_job_runner.py:705} INFO - Executor reports execution of mydag.spark_submit run_id=manual__2024-11-02T10:19:30.618008+00:00 exited with status up_for_retry for try_number 470
Configs:
with DAG(
dag_id=DAG_ID,
start_date=MYDAG_START_DATE,
schedule_interval="@daily",
catchup=AIRFLOW_CATCHUP,
default_args={
'depends_on_past': True,
"retries": 1000,
"retry_delay": duration(minutes=5),
"retry_exponential_backoff": True,
"max_retry_delay": duration(hours=1),
},
) as dag:
What you think should happen instead?
My pull request:
#48057
#48051
How to reproduce
Use configs from above. Example:
with DAG(
dag_id=DAG_ID,
start_date=MY_AIRFLOW_START_DATE,
schedule_interval="@daily",
catchup=AIRFLOW_CATCHUP,
default_args={
'depends_on_past': True,
"retries": 1000,
"retry_delay": duration(minutes=5),
"retry_exponential_backoff": True,
"max_retry_delay": duration(hours=1),
},
) as dag
Operating System
Ubuntu 22.04
Versions of Apache Airflow Providers
No response
Deployment
Official Apache Airflow Helm Chart
Deployment details
No response
Anything else?
No response
Are you willing to submit PR?
Code of Conduct