Showing posts with label unique. Show all posts
Showing posts with label unique. Show all posts

Wednesday, March 21, 2012

purge process from large table

Hi,
We have specific request here: there is a large table, around 5,000,000 rows
(3GB). It has unique clustered index, created on 6 of it's 10 columns. We
have to delete around 15% of rows every day, but in a way that table keeps
being available all the time. Deletion criteria is date (where
column_date<=getdate()). There is nonclustered index created on column_date
table. Since there is availability criteria, simple: delete <table name>
where column_date<=getdate() is out of the question because of exclusive
table lock on this table.
Any ideas?
Thanks,
PedjaBy "available" do you mean readable via a SELECT?
If you are using SQL 2005, you can take advantage of the new Read Committed
Snapshot Isolation level, which allows a SELECT to read the most
recently-committed version of a set of data, even while that data is being
modified.
If you are using SQL 2000, your readers can use WITH (NOLOCK) as an option
to the SELECT statements, putting the transaction into read uncommited. Your
readers won't have to wait for writers, but you may get inaccurate data.
Is it possible that you can split this delete operation so that it occurs
several times a day? That way it can delete fewer rows, resulting in less
blocking time.
"Pedja" wrote:

> Hi,
> We have specific request here: there is a large table, around 5,000,000 ro
ws
> (3GB). It has unique clustered index, created on 6 of it's 10 columns. We
> have to delete around 15% of rows every day, but in a way that table keeps
> being available all the time. Deletion criteria is date (where
> column_date<=getdate()). There is nonclustered index created on column_dat
e
> table. Since there is availability criteria, simple: delete <table name>
> where column_date<=getdate() is out of the question because of exclusive
> table lock on this table.
> Any ideas?
> Thanks,
> Pedja|||Mark,
We use sql server 2000. By available, I mean both, read/write operations.
NOLOCK hint won't help, because once it is grabbed by purge process, table i
s
being locked until it is completed (that is why I posted this question
initially), so neither reads nor writes are allowed during this time. My ide
a
was to split deletion process to batches of 1000 rows (set rowcount 1000),
but I wanted to hear some other ideas too.
Thanks
"Mark Williams" wrote:
> By "available" do you mean readable via a SELECT?
> If you are using SQL 2005, you can take advantage of the new Read Committe
d
> Snapshot Isolation level, which allows a SELECT to read the most
> recently-committed version of a set of data, even while that data is being
> modified.
> If you are using SQL 2000, your readers can use WITH (NOLOCK) as an option
> to the SELECT statements, putting the transaction into read uncommited. Yo
ur
> readers won't have to wait for writers, but you may get inaccurate data.
> Is it possible that you can split this delete operation so that it occurs
> several times a day? That way it can delete fewer rows, resulting in less
> blocking time.
> --
> "Pedja" wrote:
>|||Pedja,
you could split up your data into several tables, one table per day.
You can access them via a UNION ALL view. Then the purge is very fast,
you just re-create the view, which is a snap, and drop the oldest
table. There are some divantages: some of queries against the view
will work slower, and you will not be able to enforse unique (and
sometimes other) constraints just as easily|||Hi
Divide your deletion into small batches
SET ROWCOUNT 1000
WHILE 1 = 1
BEGIN
--Here your DML Statement
IF @.@.ROWCOUNT = 0
BEGIN
BREAK
END
END
SET ROWCOUNT 0
"Pedja" <Pedja@.discussions.microsoft.com> wrote in message
news:38445382-60B6-4327-A3FE-C5AD8BF88E41@.microsoft.com...
> Hi,
> We have specific request here: there is a large table, around 5,000,000
> rows
> (3GB). It has unique clustered index, created on 6 of it's 10 columns. We
> have to delete around 15% of rows every day, but in a way that table keeps
> being available all the time. Deletion criteria is date (where
> column_date<=getdate()). There is nonclustered index created on
> column_date
> table. Since there is availability criteria, simple: delete <table name>
> where column_date<=getdate() is out of the question because of exclusive
> table lock on this table.
> Any ideas?
> Thanks,
> Pedjasql

pulling unique records from this query

Hi guys, need your help! (sorry this is quite long)
I've got a table of Projects which I'm using with an asp:Repeater to display
a list of the projects. Here's the sql...
SELECT ProjectID, ProjectName, ProjectClient, DartsContact, LeadArtist,
Projects.AreaOfWork, AreaOfDoncaster, StartDate, EndDate, Running,
WorkAreas.AreaofWork AS AOWName, WorkAreas.RelatesTo AS AOWRelates FROM
Projects, WorkAreas
WHERE (NOT Running=0) AND (Projects.AreaOfWork LIKE '%' + WorkAreas.AOWCode
+ '%') AND (Deleted = 0) ORDER BY ProjectName ASC
As you can hopefully see, I'm using two tables to pull the data together.
It worked fine, until we made a change to the way the data is stored. The
Projects.AreaOfWork field now contains multiple AOWCodes seperated by a
delimiter. So now, whenever I run this query, I get more than 1 line for eac
h
project where there are multiple values in AreaOfWork. So, if I have a
project...
ProjectID, ProjectName, AreaOfWork
1, Test Proj 1, EDU
2, Test Proj 2, EDU|COM
I get 2 lines for Test Project 2, each with a unique AreaOfWork (one with
EDU, one with COM).
2 things... I need to stop it returning multiple records for the same
project when theres more than one AreaOfWork, but I also need to return the
entire AreaOfWork string, because I still need access to those.
Any help would be greatly appreciated.
Cheers
Danwhy have you chosen to store your data like that (delimited) - it
breaks with normalisation, and is the main reason your having problems.
surely it would be easier if you had a separate table like
tblProjWorkAreas(ProjectID, AreaOfWork). Is there are reason for not
doing this?|||I see your point mate. Time constraints are the main reason for this.. it's
an addition to a project that's been running for a couple of years (the
having multiples instead of one).
Is there a way to get the query to work!?
"Will" wrote:

> why have you chosen to store your data like that (delimited) - it
> breaks with normalisation, and is the main reason your having problems.
> surely it would be easier if you had a separate table like
> tblProjWorkAreas(ProjectID, AreaOfWork). Is there are reason for not
> doing this?
>|||actually, further to this... I *think* i can do half of what I want to do in
code, if I can get it to just select unique records... :)
"Will" wrote:

> why have you chosen to store your data like that (delimited) - it
> breaks with normalisation, and is the main reason your having problems.
> surely it would be easier if you had a separate table like
> tblProjWorkAreas(ProjectID, AreaOfWork). Is there are reason for not
> doing this?
>|||depends on what you want out. If we take your example where you have
EDU|COM, what output would you want - the area of work columns x2?
could you post a fuller example in terms of data from both tables, and
what you'd like your query to result in.|||Hi Dan,
Thanks for using MSDN Managed Newsgroup Support.
As Will mentioned, it is not a good idea to store your data like that.
So I want to know why you use the WorkAreas.AreaofWork and
WorkAreas.Relates in the query. If you want to unique the only Project in
this query, I think you may need to exclude the WorkAreas Table.
Also, you may provide me the result of the query now if you have 2
AreaOfWork in the Project Table.
I need to know the exactly different of these 2 records.
Sincerely,
Wei Lu
Microsoft Online Community Support
========================================
==========
When responding to posts, please "Reply to Group" via your newsreader so
that others may learn and benefit from your issue.
========================================
==========
This posting is provided "AS IS" with no warranties, and confers no rights.|||Thank you both for your help. I managed to get this to work using a differen
t
method, so no need to concern yourselves anymore :) I do appreciate that
using a third table would be a much better solution, and may consider that
for future redevlopment.
In answer to your question Wei, the reason for showing the AOW fields is
simply to show which Areas of Work a Project belongs to. WorkAreas.Relates i
s
a field that allows us to have inherited Areas of Work in the table, like so
.
aowcode aowname relates_to
1 Education null
2 Adult Ed 1
3 Preschool 1
etc.
Again, thank you both for your help today!
Cheers
Dan
"Wei Lu" wrote:

> Hi Dan,
> Thanks for using MSDN Managed Newsgroup Support.
> As Will mentioned, it is not a good idea to store your data like that.
> So I want to know why you use the WorkAreas.AreaofWork and
> WorkAreas.Relates in the query. If you want to unique the only Project in
> this query, I think you may need to exclude the WorkAreas Table.
> Also, you may provide me the result of the query now if you have 2
> AreaOfWork in the Project Table.
> I need to know the exactly different of these 2 records.
> Sincerely,
> Wei Lu
> Microsoft Online Community Support
> ========================================
==========
> When responding to posts, please "Reply to Group" via your newsreader so
> that others may learn and benefit from your issue.
> ========================================
==========
> This posting is provided "AS IS" with no warranties, and confers no rights
.
>

Tuesday, March 20, 2012

Pulling a unique field w/ join

Hi,

I'm trying to pull some data from two tables, and I need the data where a particular field is unique. Here's what I have right now:

SELECT tblTrip.*, tblTDYTravel.* FROM tblTrip
INNER JOIN tblTDYTravel ON tblTrip.TripID = tblTDYTravel.TripID
WHERE tblTrip.EmployeeID = @.EmployeeID
ORDER BY TangoID ASC

I need to get only the entries where the TangoID field is unique. Any ideas?

Thanks!What table is TangoID in?

blindman|||Originally posted by blindman
What table is TangoID in?

blindman

tblTrip|||Originally posted by Tarkon
Hi,

I'm trying to pull some data from two tables, and I need the data where a particular field is unique. Here's what I have right now:

SELECT tblTrip.*, tblTDYTravel.* FROM tblTrip
INNER JOIN tblTDYTravel ON tblTrip.TripID = tblTDYTravel.TripID
WHERE tblTrip.EmployeeID = @.EmployeeID
ORDER BY TangoID ASC

I need to get only the entries where the TangoID field is unique. Any ideas?

Thanks!

Use group by and aggregate functions.|||Originally posted by snail
Use group by and aggregate functions.

You uh, mind being a little more specific?|||This should do what you SAID you want (only unique tangoIDs):

SELECT tblTrip.*, tblTDYTravel.*
FROM tblTrip
INNER JOIN tblTDYTravel ON tblTrip.TripID = tblTDYTravel.TripID
INNER JOIN (select TangoID From TBLTrip group by TangoID Having count(*) = 1) UniqueTangoIDs
on tblTrip.TangoID = UniqueTangoIDs.TangoID
WHERE tblTrip.EmployeeID = @.EmployeeID
ORDER BY tblTrip.TangoID ASC

..whether this is what you MEAN you want is another story.

blindman