Repository navigation
Conversation
a4127a2 to
5499321
Compare
601e5bb to
fc0db1a
Compare
|
After the last test fix I did there, I'm thinking what would the correct default behavior for this feature should be? Currently it's default True, but if we want to normally just skip the TI, then the default should be False. |
|
@RNHTTR Checking in here, have you gotten a chance to look at these changes? |
There was a problem hiding this comment.
Should this have a specific exception handling for AttributeError in the case where the operator doesn't have an on_kill method defined? In this case, I think we'd want to self.log.info("on_kill method does not exist for task %s: %s", task_instance, e).
What do you think?
There was a problem hiding this comment.
Or maybe something like...
if hasattr(task, 'on_kill'):
try:
...
else:
self.log.warn("Task %s was configured to execute its on_kill method after a DAG run timeout, but it does not have the method 'on_kill' method defined", task_instance)There was a problem hiding this comment.
If there can be a case of an operator without an on_kill, then definitely we should check for that. I think using hasattr to check for it before getting the task and using on_kill is better to prevent the AttributeError. I will make this change.
9d373b5 to
316888c
Compare
|
I had a conflict and made a mistake when pushing the commits again. After getting the commits pushed, the build is failing with the error: |
9b14cac to
8250b6b
Compare
There was a problem hiding this comment.
Is there any context/previous discussion about this?
The reason I ask is that I'm wary of adding yet more options of ways things can run
There was a problem hiding this comment.
@RNHTTR can help clarify, but basically this change would allow tasks to run on_kill when a dagrun reaches timeout. The use case given in the issue is that "Some users would like for externally running workloads (e.g. snowflake, emr, bigquery, etc etc) to stop executing when a DAG run times out." . With this flag the user will have more control on what behavior they want when there's a timeout.
There was a problem hiding this comment.
This all stems from the fairly controversial idea of setting running tasks to skipped if a DAG run reaches its timeout.
There are a couple paths forward:
- Change the behavior after a
dagrun_timeoutis reached to mark running TIs as failed. - Proceed with this PR
- Encourage users to use an on_skipped_callback for external systems that need to be shutdown when a task moves to the skipped state after a DAG run timeout.
There was a problem hiding this comment.
IMO this would be a good change. It's somewhat counterintuitive that tasks will be marked as skipped when a Dagrun times out, and enabling on_kill on Dagrun timeout is intuitive.
There was a problem hiding this comment.
@ashb Any thoughts on what's the best path forward?
Also, I mentioned in a comment here that there has been a significant change to the DAG properties. If we want to use this new call_on_kill_on_dagrun_timeout, where should it go now?
654931a to
b8148b1
Compare
1a20120 to
06fd3b4
Compare
Co-authored-by: Ryan Hatter <25823361+RNHTTR@users.noreply.github.com>
06fd3b4 to
5e4df97
Compare
|
Looks like there has been an extensive change that was merged recently here that changed how the DAG properties are defined. I'm going through the changes in the PR, but would the changes from here go in Airflow's core rather than in the Task SDK? Trying to fix the merge conflict and a bit confused in where the properties go now. |
|
@MRLab12 I think they definitely need to go into the Task SDK, because that's what users will interface with. Because the models inherit from their associated SDK objects (example), I don't think you need to duplicate them there unless you'll be explicitly writing code at the model level, which I don't think is necessary. |
|
This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 5 days if no further activity occurs. Thank you for your contributions. |
This PR addresses issue #41036 by adding support for the
on_killedcallback on tasks that are still running when a DAG run reaches its timeout.Key changes:
call_on_kill_on_dagrun_timeoutto control whether tasks should be killed when a DAG run times out (default: True)._schedule_dag_runto call taskon_killifcall_on_kill_on_dagrun_timeoutis enabled.closes: #41036
Testing: