Skip to main content
Member
September 16, 2026
Solved

Joining Data From Multiple Sources

  • September 16, 2026
  • 1 reply
  • 40 views

Can you use dataframes to join data from multiple sources or join multiple dataframes into one? For example, a mapping table in an external database that we’d like to join to a custom table in the OS application database, before doing additional processing (grouping, aggregation, etc) on the joined data.

Best answer by JackLacava

 

 

Transcript

K goldbach asks whether you can use DataFrames to join data from multiple sources or joining multiple DataFrames into one. At the moment with the current API, you can append new DataFrame to an existing one. The append method, as long as their columns are the same. That works fine, but if you're talking about the sort of thing that data set objects do, like mapping between one field in one record set and the other record set, that sort of thing, we don't currently have it.

We are thinking about ways of possibly enabling this in the future, but at the moment we don't have it. So the best approach is probably to do your join with an actual SQL query that performs a join, and execute it through a DataFrame. So you get a DataFrame back because it's super quick. Because of the speed gains, it's still likely to be faster than going through the whole data set of data set operation.

I know the interface is going to be different from a from a coding perspective. The feeling might be different. But chances are the, particularly if they're working with large data sets, the speed improvement will be worth the time you spend changing the code.

1 reply

Sage
September 22, 2026

 

 

Transcript

K goldbach asks whether you can use DataFrames to join data from multiple sources or joining multiple DataFrames into one. At the moment with the current API, you can append new DataFrame to an existing one. The append method, as long as their columns are the same. That works fine, but if you're talking about the sort of thing that data set objects do, like mapping between one field in one record set and the other record set, that sort of thing, we don't currently have it.

We are thinking about ways of possibly enabling this in the future, but at the moment we don't have it. So the best approach is probably to do your join with an actual SQL query that performs a join, and execute it through a DataFrame. So you get a DataFrame back because it's super quick. Because of the speed gains, it's still likely to be faster than going through the whole data set of data set operation.

I know the interface is going to be different from a from a coding perspective. The feeling might be different. But chances are the, particularly if they're working with large data sets, the speed improvement will be worth the time you spend changing the code.