Repository navigation
Status of testing Providers that were prepared on November 24, 2022 #27894
Description
Activity
- addedkind:metaHigh-level information important to the communityHigh-level information important to the community
on Nov 24, 2022 #27724 works as expected! Finally Trino inserts are 100% working
Reacted by Jarek PotiukReacted by Jarek Potiuk#27868 works, #27854 unfortunately not - results are double wrapped:
For example for
select * from default.my_***_table, parameters: Nonewe get:scalar_results=True,results=[Row(id=1, v='test 1'), Row(id=2, v='test 2')]. And then code:if scalar_results: list_results: list[Any] = [results] else: list_results = resultswraps this list into another list. I think that logic should be changed a bit - right now we're collecting results for all SQL statements into a single list although they could have different schemas.
Let me debug it and open another PR
Reacted by Jarek Potiukwraps this list into another list. I think that logic should be changed a bit - right now we're collecting results for all SQL statements into a single list although they could have different schemas.
Yeah. that part is a bit not clear about the intentions (or maybe I misunderstood it). Woudl be great if you have a PR indeed.
Ah I think I see where I made wrong assumption @alexott . looking at it
In Databricks SQL operator (and I believe in others as well), there was following strategy: always return only last result - previous results were always discarded. Primary reason for this was following:
- When you have multiple SQL statements, first one usually create table, inserts, etc. And only when you have select as the last statement, then you get results. This matches the logic of the SQL's
BATCHstatement - When you have multiple SQL statements their result may have different schema, but results will be processed only according to the latest schema, not schemas for corresponding result sets
We may need to think a bit about it - should we return results for each of the statements, or not. If yes, then we need to return pairs of description + results for each SQL statement, instead of using only the latest statement
- When you have multiple SQL statements, first one usually create table, inserts, etc. And only when you have select as the last statement, then you get results. This matches the logic of the SQL's
#27276 works as expected. System tests using those operators are all passing.
Reacted by Jarek PotiukIn Databricks SQL operator (and I believe in others as well), there was following strategy: always return only last result - previous results were always discarded. Primary reason for this was following:
- When you have multiple SQL statements, first one usually create table, inserts, etc. And only when you have select as the last statement, then you get results. This matches the logic of the SQL's
BATCHstatement - When you have multiple SQL statements their result may have different schema, but results will be processed only according to the latest schema, not schemas for corresponding result sets
We may need to think a bit about it - should we return results for each of the statements, or not. If yes, then we need to return pairs of description + results for each SQL statement, instead of using only the latest statement
Yes - I noticed that too now. With two caveats:
depends on the oprator what is the default (no problem)- nope - all of the use True. Fine- it behaves differently when there is an "sql" passed and return_last is true -> then instead of one-element result array it returns the results
It is surprisingly difficult to unwind teh original convoluted behaviour :)
- When you have multiple SQL statements, first one usually create table, inserts, etc. And only when you have select as the last statement, then you get results. This matches the logic of the SQL's
Opened #27897 - tested all file formats
#26986 Works as expected. I am able to view the DataprocLink irrespective of Job status
Reacted by Jarek Potiuk#26374 tested, work as described
Reacted by Jarek PotiukI will have to cancel the vote due to a bug found still in common.sql. Much more complete and comprehensive (with many tests added) fix is coming #27912
I will carry-over the status of testing for the issues that have been already tested (thanks to everyone who helped with testing! ) so hopefully the last round of testing can focus on sql only.
I have a kind request for all the contributors to the latest provider packages release.
Could you please help us to test the RC versions of the providers?
Let us know in the comment, whether the issue is addressed.
Those are providers that require testing as there were some substantial changes introduced:
Provider amazon: 6.2.0rc2
RedshiftResumeClusterOperatorandRedshiftPauseClusterOperator(#27276): @syedahsnProvider asana: 2.1.0rc2
Provider common.sql: 1.3.1rc2
Provider databricks: 4.0.0rc2
Provider exasol: 4.1.1rc2
Provider google: 8.6.0rc2
Provider jdbc: 3.3.0rc2
Provider mysql: 3.4.0rc2
Provider neo4j: 3.2.1rc2
Provider presto: 4.2.0rc2
Provider slack: 7.1.0rc2
Provider snowflake: 4.0.1rc2
Provider trino: 4.3.0rc2
The guidelines on how to test providers can be found in
Verify providers by contributors