Showing posts with label relation. Show all posts
Showing posts with label relation. Show all posts

Friday, March 9, 2012

Fastest way of updating a row

In relation to my last post, I have a question for the SQL-gurus.

I need to update 70k records, and mark all those updated in a special
column for further processing by another system.

So, if the record was

Key1, foo, foo, ""

it needs to become

Key1, fap, fap, "U"

iff and only iff the datavalues are actually different (as above, foo
becomes fap),

otherwise it must become

Key1, foo,foo, ""

Is it quicker to :
1) get the row of the destination table, inspect all values
programatically, and determine IF an update query is needed

OR

2) just do a update on all rows, but adding
and (field1 <> value1 or field2<>value2) to the update query

that is
update myTable
set
field1 = "foo"
markField="u"
where key="mykey" and (field1 <> foo)

The first one will not generate new update queries if the record has
not changed, on account of doing a select, whereas the second version
always runs an update, but some of them will not affect any lines.

Will I need a full index on the second version?

Thanks in advance,
Asger Henriksen[posted and mailed, vnligen svara i nys]

Asger Jensen (akj@.tmnet.dk) writes:
> Is it quicker to :
> 1) get the row of the destination table, inspect all values
> programatically, and determine IF an update query is needed
> OR
> 2) just do a update on all rows, but adding
> and (field1 <> value1 or field2<>value2) to the update query
> that is
> update myTable
> set
> field1 = "foo"
> markField="u"
> where key="mykey" and (field1 <> foo)

I'm not sure that I follow, but it sounds to me that the in first
approach you would retrieve rows one by one.

In any case, the second approach leaves all the jub to the computer,
and there is a reason why we have computers, isn't there? :-)

The only catch is that with too many rows in the table there can be
a strain on the transaction log. But with only 70000 rows, this is
not worth worrying about.

Obviously there query will run faster if there is a clustered index
on the column "key".

--
Erland Sommarskog, SQL Server MVP, sommar@.algonet.se

Books Online for SQL Server SP3 at
http://www.microsoft.com/sql/techin.../2000/books.asp|||> I'm not sure that I follow, but it sounds to me that the in first
> approach you would retrieve rows one by one.

Yes, thats what it was. Nasty.
> In any case, the second approach leaves all the jub to the computer,
> and there is a reason why we have computers, isn't there? :-)
:-), yes, only in my scenario it took soooo long.

> The only catch is that with too many rows in the table there can be
> a strain on the transaction log. But with only 70000 rows, this is
> not worth worrying about.

ok,

> Obviously there query will run faster if there is a clustered index
> on the column "key".
There was, actually, but during the 70000 records it became slower and
slower. I solved it by adding a dynamically generated index on ALL
fields, not just "key", this sped things up considerably, so I went
from 1 hour to 10 minutes.

I guess it is due to the server needing to do a lookup/record on all
value fields to see if they have changed, and this can be done more
efficiently with a n index.

Regards
Asger

Wednesday, March 7, 2012

Fast loading relational data

I am searching for a way to fast load relation data. I know how to load data fast but how can i store relation data fast.
For example :
Table1 ( tabel1Id int identity , name varchar(255) )
Table2 ( tabel2Id int identity , table2Id , name varchar(255))
When i insert 50 records into Table1 i can't get the 50 identity fields back, to insert the related data into Table2.
I think one of the solutions could be returning a selection of Table1 joined with syslockinfo, but i have no idea how to do it.
Does anyone have an idea?

Add uniqueidentifier column to your master table.

When inserting, generate the value by newid() function and insert it alongside your 50 records. Then you can easily find all the records inserted.

|||Thank you for your reply.
I also thought of that. But i also thought that it must be possible to join with syslockinfo.
|||

I would suggest that you correlate the values in the insert. So, say you have to insert an invoice and line items.

invoiceNumber
000000333
000000334

invoiceNumber product amount
000000333 widget 10
000000334 wadget 11
then the invoice insert is easy, and the line item insert is something like:

insert into invoiceLineItem (invoiceId, product, amount)
select invoiceId, product, amount
from invoice
join (select '000000333' as invoiceNumber,'widget' as product,10 as amount
union all
select '000000334','wadget',11
) as newInvoiceLines
on invoice.invoiceNumber = newInvoiceLines.invoiceNumber

Of course exactly how you get the bold stuff will depend on the format of your data (like a spreadsheet, or from a user interface, XML, etc. In reality I would have expected you to have to get a productId as well.