Issue Description
I am incrementally slinging data to a large Oracle table via sqlldr from SQL Server. Here is an example command:
sling run --src-conn CDW_MSSQL --src-stream dbo.BillingTransactionFact --tgt-conn ORABOODLE --tgt-object BillingTransactionFact --tgt-conn ORABOODLE --tgt-object BillingTransactionFact --mode incremental --update-key BillingTransactionKey --limit 250000 -d
The table rows have a lot of data, which requires a lot of space to manage when ordering is required. For this table, I must move in very small increments to avoid the following error:
mssql: Could not allocate a new page for database 'TEMPDB' because the 'DEFAULT' filegroup is full due to lack of storage space or database files reaching the maximum allowed size. Note that UNLIMITED files are still limited to 16TB. Create the necessary space by dropping objects in the filegroup, adding additional files to the filegroup, or setting autogrowth on for existing files in the filegroup.
- Description of the issue:
The command spawns a query that is logged like this:
2024-11-21 18:59:16 INF reading from source database
2024-11-21 18:59:16 DBG select top 250000 * from "dbo"."BillingTransactionFact" where "BillingTransactionKey" > 12249997 order by "BillingTransactionKey" asc
This query, after considerable processing time, produces the TEMPDB error above.
The query produces the same error outside sling.
The following equivalent query, on the other hand, is both more performant and uses much less TEMPDB space:
with ids as (
select top 400000 "BillingTransactionKey"
from "dbo"."BillingTransactionFact"
where "BillingTransactionKey" > 12249997
order by "BillingTransactionKey" asc
) select *
from "dbo"."BillingTransactionFact" btf
inner join ids i
on btf."BillingTransactionKey" = i."BillingTransactionKey";
Could sling use logic like this to read from a SQL Server source database?
I think this optimization could help a lot.
-
Sling version (sling --version):
Version: 1.2.22
-
Operating System (linux, mac, windows):
Windows
-
Replication Configuration:
N/A
-
Log Output (please run command with -d):
See above.
Issue Description
I am incrementally slinging data to a large Oracle table via sqlldr from SQL Server. Here is an example command:
The table rows have a lot of data, which requires a lot of space to manage when ordering is required. For this table, I must move in very small increments to avoid the following error:
The command spawns a query that is logged like this:
This query, after considerable processing time, produces the TEMPDB error above.
The query produces the same error outside sling.
The following equivalent query, on the other hand, is both more performant and uses much less TEMPDB space:
Could sling use logic like this to read from a SQL Server source database?
I think this optimization could help a lot.
Sling version (
sling --version):Version: 1.2.22
Operating System (
linux,mac,windows):Windows
Replication Configuration:
N/A
Log Output (please run command with
-d):See above.