Archive for the 'oracle' Category



Quantifying Hugepages Memory Savings with Oracle Database 11g

In my recent post about physical memory consumed by page tables when hugepages are not in use I showed an example of 500 dedicated connections to an Oracle Database 11g instance with 8000M SGA consuming roughly 7 gigabytes of physical memory just for page tables. A reader emailed me to point out that it would be informational to show page table consumption with hugepages employed. How right.

The following script was running while I invoked sqlplus 500 times to connect to the same instance discussed in this post. As the script shows, page table cost peaked at roughly 265 MB compared to the 7 GB lost to page tables in the non-hugepages case.

I’m scratching my head and thinking of what else could possibly be said on the matter…

$ while true
> do
> ps -ef | grep oracletest | wc -l
> grep PageTables /proc/meminfo
> sleep 30
> done
1
PageTables:      23484 kB
501
PageTables:     264432 kB
501
PageTables:     264900 kB
501
PageTables:     265400 kB
501
PageTables:     265944 kB
130
PageTables:      83112 kB
120
PageTables:      78672 kB
110
PageTables:      74188 kB
100
PageTables:      69712 kB
90
PageTables:      65264 kB
80
PageTables:      60804 kB
70
PageTables:      56332 kB
60
PageTables:      51872 kB
50
PageTables:      47376 kB
40
PageTables:      42848 kB
30
PageTables:      37904 kB
20
PageTables:      32976 kB
10
PageTables:      28024 kB
1
PageTables:      23532 kB
1
PageTables:      23528 kB

Little Things Doth Crabby Make – Part X. Posts About Linux Hugepages Makes Some Crabby It Seems. Also, Words About Sizing Hugepages.

I received a few pieces of (not)fan-mail about my latest post in the Crabby Series. One reader took offense at the fact that I bother to blog about hugepages because, in his words:

…you insult the intelligence of your readers. You know full well everyone uses hugepages

Is that why Metalink Note 749851.1 goes to the trouble of advising DBAs that the default database setup from Database Configuration Assistant (DBCA) configures Automatic Memory Management which does not use hugepages?

I assure you, not everyone uses hugepages and part of that is because it can be difficult to set it up if you have several databases—especially if your databases have a mix of heavy PGA usage and heavy SGA usages. Also, if your calculations are off and there are insufficient hugepages to cover the SGA, Oracle will go ahead and allocate with a shmget() that doesn’t pass in SHM_HUGETLB. The effect of that little twist is you’ll be “missing” the memory that was carved out for hugepages and the SGA will reside in other non-hugepages memory. So, for instance, if you calculate your SGA to be 1GB and you allocate 513 (1GB + 1 page for wiggle room) but your SGA turns out to be 1073758208 (1GB + 16KB), you’ll get a non-hugepages SGA and eventually there will be roughly 2GB tied up. I think it is an important topic.

Metalink 401749.1
Oracle Support offers a script to assist DBAs in calculating hugepages requirement. With all your instances up, run the script and it will calculate a setting for you. The note is entitled Shell Script to Calculate Values Recommended HugePages / HugeTLB Configuration.

There is a small nit regarding this note ( the procedure it involves actually).  In order for the script to give you a recommendation, you have to revert from AMM first, then do a boot of your instances with MMM so it can peek what SysV IPC segments are being allocated for the instances. So, it’s a multi-step process. I suppose with a lot of extra thought the same thing could be calculated by tallying up all the “granule files” found in /dev/shm under AMM, but no matter. This is fairly simple.

Let’s look at my system. Here’s what we’ll see:

First, we’ll see how large the SGA really is.

Next, we’ll see how large of an IPC segment the instance called for. In my case it is about 37MB larger than the actual SGA. That’s fine.

Finally we’ll see the output of the hugepages_settings.sh script to see what it advises.

SQL*Plus: Release 11.X.0.X.0 Production on Mon Jul 27 13:35:41 2009

Copyright (c) 1982, 2009, Oracle.  All rights reserved.

Connected to:
Oracle Database 11g Enterprise Edition Release 11.X.0.X.0 - 64bit Production
With the Partitioning, OLAP, Data Mining and Real Application Testing options

SQL> show sga

Total System Global Area 8351150080 bytes
Fixed Size                  2214808 bytes
Variable Size            1543505000 bytes
Database Buffers         6777995264 bytes
Redo Buffers               27435008 bytes
SQL> Disconnected from Oracle Database 11g Enterprise Edition Release 11.X.0.X.0 - 64bit Production
With the Partitioning, OLAP, Data Mining and Real Application Testing options
$ ipcs -m

------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status
0x522d5fd4 327681     oracle    660        8390705152 53                      

$ sh ./hugepages_setting.sh
Recommended setting: vm.nr_hugepages = 4003
$ grep Huge /proc/meminfo
HugePages_Total:  5000
HugePages_Free:    999
HugePages_Rsvd:      0
Hugepagesize:     2048 kB

So it looks like the script is accurate and even allows a little wiggle room. That’s good. I think this script (being helpful) combined with a healthy fear of the nastiness in large SGA+large dedicated connection deployments (without hugepages) should get us all one step closer to insisting on hugepages backing for our SGAs.

Little Things Doth Crabby Make – Part IX. Sometimes You Have To Really, Really Want Your Hugepages Support For Oracle Database 11g.

Recently I had someone ask me in email why I bother posting installments on my Little Things Doth Crabby Make series. I responded by saying I think it is valuable to IT professionals to know they are not alone when confronted by something that makes little sense, or makes them crabby if that be the case. It’s all about the Wayward Googler(tm).

Well, Wayward Googler, it’s coming on thick.

Using Memory and Then Allocating HugePages (Or Die Trying)
I purposefully booted my system with no hugepages allocated in /etc/sysctl.conf (vm.nr_hugepages = 0). I then booted an Oracle Database 11g instance with sga_target set to 8000M. Next, I fired off 500 dedicated connections using the following goofy stuff:


$ cat doit
cnt=0
until [ $cnt -eq 500 ]
do
   sqlplus rw/rw @foo.sql &
   (( cnt = $cnt + 1 ))
done

wait

$ cat foo.sql
HOST sleep 120
exit;

The script ran in a matter of moments since I’m using a Xeon 5500 (Nehalem) based dual-socket server running Linux with a 2.6 kernel. Yes, these processors are really, really fast. But that, of course, isn’t what made me crabby.

Directly before I invoked the script, that fired off my 500 dedicated connections,  I executed a script that intermittently peeked at how much memory is being wasted on page tables. Remember, without hugepages (hugetlb) backed IPC Shared Memory for the SGA there will be page table overhead for every connection to the instance. The size of the SGA and the number of dedicated connections compounds to consume potentially significant amounts of memory. Although that is also not what made me crabby, let’s look at what 500 dedicated sessions attaching to an 8000 MB SGA looks like as the user count ramps up:


$ while true
> do
> grep PageTables /proc/meminfo
> sleep 10
> done

PageTables:       3764 kB
PageTables:       4696 kB
PageTables:      65848 kB
PageTables:     176956 kB
PageTables:     287616 kB
PageTables:     366540 kB
PageTables:     478224 kB
PageTables:     588424 kB
PageTables:     699832 kB
PageTables:     792356 kB
PageTables:     802468 kB
PageTables:     834004 kB
PageTables:     851980 kB
PageTables:     835432 kB
PageTables:     834948 kB
PageTables:     835052 kB
PageTables:    1463260 kB
PageTables:    2072864 kB
PageTables:    2679572 kB
PageTables:    3283456 kB
PageTables:    3892628 kB
PageTables:    4496868 kB
PageTables:    5100908 kB
PageTables:    6846256 kB
PageTables:    6866820 kB
PageTables:    6829388 kB
PageTables:    6874752 kB
PageTables:    6879360 kB
PageTables:    6883076 kB
PageTables:    6895244 kB
PageTables:    6901528 kB
PageTables:    6917256 kB
PageTables:    6927984 kB
PageTables:    6999196 kB
PageTables:    6999472 kB
PageTables:    7000048 kB
PageTables:    7088160 kB
PageTables:    7087960 kB
PageTables:    7088812 kB
PageTables:    7132804 kB
PageTables:    7121120 kB

Got Spare Memory? Good, Don’t Use Hugepages
Uh, just short of 7 GB of physical memory lost to page tables! That’s ugly, but that’s not what made me crabby. Before I forget, did I mention that it is a really good idea to back your SGA with hugepages if you are running a lot of dedicated connections and have a large SGA?

So, What Did Make Him Crabby Anyway?
Wasting all that physical memory with page tables was just part of some analysis I’m doing. I never aim to waste memory (nor processor cycles for TLB misses) like that. So, I shut my Oracle Database 11g instance down in order to implement hugepages and move on. This is where I started getting crabby.

The first thing I did was verify there were, in fact, no allocated hugepages. Next, I checked to see if I had enough free memory to mess with. In this case I had most of the 16GB physical memory free. So, I tried to allocate 6200 2MB hugepages by echoing the token into /proc.  Finally, I checked to make sure I was granted the hugepages I requested…Irk. Now that, made me crabby. Instead of 6200 I was given what appears to be some random number someone pulled out of the clothes hamper—604 hugepages:

# grep HugePages /proc/meminfo
HugePages_Total:     0
HugePages_Free:      0
HugePages_Rsvd:      0
# free
             total       used       free     shared    buffers     cached
Mem:      16427876     422408   16005468          0      24104     209060
-/+ buffers/cache:     189244   16238632
Swap:      2097016      29836    2067180
# echo 6200 > /proc/sys/vm/nr_hugepages
# grep HugePages /proc/meminfo
HugePages_Total:   604
HugePages_Free:    604
HugePages_Rsvd:      0

So, I then checked to see what free memory looked like:

# free
             total       used       free     shared    buffers     cached
Mem:      16427876    1670400   14757476          0      27040     207924
-/+ buffers/cache:    1435436   14992440
Swap:      2097016      29696    2067320

Clearly I was granted that oddball 604 hugepages I didn’t ask for. Maybe I’m supposed to just take what I’m given and be happy?

Please Sir, May I Have Some More?

I thought, perhaps the system just didn’t hear me clearly. So, without changing anything I just belligerently repeated my command and found that doing so increased my allocated hugepages by a whopping 2:

# echo 6200 > /proc/sys/vm/nr_hugepages
# grep HugePages /proc/meminfo
HugePages_Total:   608
HugePages_Free:    608
HugePages_Rsvd:      0

I began to wonder if there was some reason 6200 was throwing the system a curve-ball. Here’s what happened when I lowered my expectations by requesting 3100:

# echo 3100 > /proc/sys/vm/nr_hugepages;grep HugePages /proc/meminfo
HugePages_Total:   610
HugePages_Free:    610
HugePages_Rsvd:      0

Great. I began to wonder how long I could continually whack my head against the wall picking up little bits and pieces of hugepages along the way. So, I scripted 1000 consecutive requests for hugepages. I thought, perhaps, it was necessary to really, really want those hugepages:

# cnt=0;until [ $cnt -eq 1000 ]
> do
> echo 6200 > /proc/sys/vm/nr_hugepages
> (( cnt = $cnt + 1 ))
> done
# grep HugePages /proc/meminfo
HugePages_Total:  5502
HugePages_Free:   5502
HugePages_Rsvd:      0

Brilliant! Somewhere along the way the system decided to start doling out more than those piddly 2-page allocations in response to my request for 6200, otherwise I would have exited this loop with 2,610 hugepages. Instead, I exited the loop with 5502.

Well, since some is good, more must be better. I decided to run that stupid loop again just to see if I could pick up any more crumbs:

# cnt=0;until [ $cnt -eq 1000 ]; do echo 6200 > /proc/sys/vm/nr_hugepages; (( cnt = $cnt + 1 )); done
# grep PageTables /proc/meminfo
PageTables:       7472 kB
# grep '^Hu' /proc/meminfo
HugePages_Total:  5742
HugePages_Free:   5742
HugePages_Rsvd:      0
Hugepagesize:     2048 kB

That makes me crabby.

Summary:
We should all do ourselves a favor and make sure we boot our servers with sufficient hugepages to cover our SGA(s). And, of course, you don’t get hugepages if you use Automatic Memory Management.

Little Things Doth Crabby Make – Part VIII. Hugepage Support for Oracle Database 11g Sometimes Means Using The ipcrm Command. Ugh.

Not that anyone should care about the things that make me crabby, but…here comes another brief post in my Little Things Doth Crabby Make series.

In the following box, you’ll see how I was just simply trying to remove a wee bit of detritus, specifically a segment of SysV IPC shared memory. So, here’s how this all transpired:

  • I used the ipcs command to get the shmid (262145).
  • I then fat-fingered a typo and tried to remove shmid 262146
  • Having realized what I did I immediately satisfied my curious morbidity, er, I mean morbid curiosity and checked the return code from the ipcs command. Oddly ipcs reported that it was perfectly happy to not remove a non-existent segment. But that’s not entirely what made me crabby.
  • I then issued the command without the typo.
  • Next (as fast as I could type) I checked to see what segments remained. That’s where I started to get even crabbier.
  • Since it seemed I was facing some odd stubbornness, I decided to issue the “old school” style command (i.e., the shm argument/option pair). The command failed but since that style of command is documented as deprecated I simply thought I was getting deprecated functionality.
  • Finally, I ran ipcs once gain to find that the segment was gone.
# ipcs -m

------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status
0x00000000 0          root      600        189        1          dest
0x522d5fd4 262145     oracle    660        6444548096 14                      

# ipcrm -m 262146
# echo $?
0
# ipcrm -m 262145
# ipcs -m

------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status
0x00000000 0          root      600        189        1          dest
0x00000000 262145     oracle    660        6444548096 14         dest 

# ipcrm shm 262145
cannot remove id 262145 (Invalid argument)
# ipcs -m

------ Shared Memory Segments --------
key        shmid      owner      perms      bytes      nattch     status
0x00000000 0          root      600        189        1          dest         

Forget the oddball situation with the return code for a moment. What I just discovered is that the work behind the ipcrm command that clears down the memory is asynchronous functionality. You all may have known that for all time, but I didn’t. Or, at least I don’t remember forgetting that fact if I did know it at one time.

It turns out the very first ipcrm –m 262145 was in the process of succeeding. That’s why my deprecated command usage was answered truthfully with EINVAL. The segment was gone, I was just being impatient.

Disclaimer
I reserve the right to remain crabby about the success code returned after my failed attempt to remove a segment that didn’t exist.

Hold it. Why would anyone care about SysV IPC shared memory where the Linux ports of Oracle Database 11g are concerned? After all, Automatic Memory Management is implemented via memory mapped files.

Summary
Don’t script against the return code of the Linux ipcrm command. It might make you crabby.

Oracle Exadata Storage Server Technical Deep Dive Series – Part II: Requires a Citrix CODEC.

Several people have pointed out that Part II in my Oracle Exadata Storage Server Technical Deep Dive Series would not play back on their computers for lack of a codec.  That is true and I didn’t know that when I tested the uploaded version because I have that particular codec installed on my system. The required bits are produced by Citrix Online. I have updated my webcast index page with more information for obtaining the required codec.

This issue only related to Part II in the series.

For what it’s worth, the codec is easily uninstalled through control panel->Add Remove Software.

Oracle Exadata Storage Server Architecture: Impossible To Back Up?

Too Large To Back Up?
If it is possible to back up a large data warehouse at rates of over 11 TB/h for full backup and more than 100 TB/h for incremental backups, maybe not!

The Maximum Availability Architecture (MAA) team has just published a paper covering tape backup of the HP Oracle Database Machine. The conclusion reads:

With the Exadata Storage Server, Oracle provides an architecture that allows customers with large databases to scale their tape backup to any desired peformance level. The number and connectivity of media servers, and the number and speed of tape drives will define the performance limit of backup, not the Database Machine. With two media servers, effective full backup rates from 11.2 TB/hour and effective incremental backup rate of over 104 TB/hour were achieved.

The title of the paper is Tape Backup Performance and Best Practices for Exadata Storage and the HP Oracle Database Machine and the paper can be accessed at the following link:

Tape Backup Performance and Best Practices for Exadata Storage and the HP Oracle Database Machine

Oracle Database File System (DBFS) on Exadata Storage Server. Hidden Content?

A colleague of mine in Oracle’s Real-World Performance Group just pointed out to me that the link (on my Papers, Webcasts, etc page) to the archived webcast of Part IV in my Oracle Exadata Storage Server Technical Deep Dive Series was stale. Actually, the problem turns out that I mistakenly set the file to expire after a fixed number of downloads. I didn’t think it would get downloaded 500 times but it seems I was wrong.

I just fixed it,  so if you’ve tried to get Part IV and hit this problem also, please give it a go now. For those of you who don’t know about this series, please visit my Papers, Webcasts, etc page where I have posted a description of each archived webcast.

Where’s Part III?

I am still on the hook for doing Part III again to recover from the loss of the IOUG recording. Part III goes into the aspect of Exadata architecture known as the division of work. Understanding the division of work is important for folks trying to decide what sort of configuration they’d have to assemble to match the performance of an Exadata deployment (e.g., HP Oracle Database Machine). I’ve been very interrupt-driven lately so I have not been able to do the Part III over again. I’ll post a blog entry when it is available.

Aren’t Customers Choosing Oracle Database Machine?

This is just a quick blog entry to point to the first heavily customer-focused news release about the Oracle Database Machine (based on Oracle Exadata Storage Server).

Here is the link:

Customers are Choosing the Oracle Database Machine

Recorded Webcast Available: Exadata Storage Server Technical Deep Dive – Part IV.

This is just a quick blog entry to point out that I updated my “Papers, etc” section with a link to the recorded Exadata Storage Server Technical Deep Dive – Part IV webcast.

Oracle-Enhancing Solaris Features. Memory Lane.

Reinventing Inventions. Deja Vu.

My old friend Glenn Fawcett sent me a link to a list of historical key technological Solaris platform enhancements for Oracle. After thanking him for that I (no surprise) felt compelled to point out which items on that list had been implemented in Sequent DYNIX/ptx on average 3 years prior to being implemented in Solaris. 🙂  Although neither of us spoke the words, but were likely thinking nonetheless, it’s interesting how many of the items on the list emerged in Linux on average several years after the Solaris rendition hit the streets.

Glenn successfully completed his ex-Sequent 12-step program many years ago. I was not as successful it seems.

Staging Data For ETL/ELT? Flat Files Appear Magically! No, Load Time Starts With Transfer Time.

In my recent post entitled Something to Ponder? What Sort of Powerful Offering Could a Filesystem in Userspace Be?, I threw out what may have seemed to be a totally hypothetical Teaser Post™. However, as our HP Oracle Exadata Storage Server and HP Oracle Database Machine customers know based on their latest software upgrade, a FUSE-based Oracle-backed file system is a reality. It is called Oracle Database File System (DBFS). DBFS is one of the corner stones of data loading infrastructure in the HP Oracle Database Machine environment. For the time being it is a staging area for flat files to be accessed as external tables. The “back-end”, as it were, is not totally new. See, DBFS is built upon Oracle SecureFiles. FUSE is the presentation layer that makes for mountable file systems. Mixing FUSE with DBFS results in a distributed, coherent file system scalable due to Real Application Clusters. This is a file system that is completely managed by Database Administrators.

So, I’m sure some folks’ eyes are rolling back in their head wondering why we need YAFS (Yet Another File System). Well, as time progresses I think Oracle enthusiasts will come to see just how feature rich something like DBFS really is.

If it performs, is feature rich and incremental to existing technology, it sounds awfully good to me!

I’ll be discussing DBFS in Oracle Exadata Technical Deep Dive – Part IV session tomorrow.

Here is a quick snippet of DBFS in action. In the following box you’ll see a DBFS mount of type FUSE on /data and the listing of a file called all_card_trans.ul


$ mount | grep fuse
dbfs on /data type fuse (rw,nosuid,nodev,max_read=1048576,default_permissions,allow_other,user=oracle)
$ pwd
/data/FS1/stage1
$ ls -l all_card_trans.ul
-rw-r--r-- 1 oracle dba 34034910300 Jun 15 15:30 all_card_trans.ul

In the next box you’ll see ssh jumping to 4 of the servers in an HP Oracle Database Machine to list the contents of the DBFS file system and md5sum output to validate that it is the same file.


$ for n in r1 r2 r3 r4
> do
> ssh $n md5sum `pwd`/all_card_trans.ul &
> done
[5] 3943
[6] 3945
[7] 3946
[8] 3947
$ 1adbff1a36a42253c453c22dd031b48b  /data/FS1/stage1/all_card_trans.ul
1adbff1a36a42253c453c22dd031b48b  /data/FS1/stage1/all_card_trans.ul
1adbff1a36a42253c453c22dd031b48b  /data/FS1/stage1/all_card_trans.ul
1adbff1a36a42253c453c22dd031b48b  /data/FS1/stage1/all_card_trans.ul
[5]   Done                    ssh $n md5sum `pwd`/all_card_trans.ul
[6]   Done                    ssh $n md5sum `pwd`/all_card_trans.ul
[7]   Done                    ssh $n md5sum `pwd`/all_card_trans.ul
[8]   Done                    ssh $n md5sum `pwd`/all_card_trans.ul

In the next box you’ll see concurrent multi-node throughput. I’ll use one dd process on each of 4 servers in the HP Oracle Database Machine each sequentially reading the contents of the same DBFS-based file and achieving 876 MB/s aggregate throughput. And, no, there is no cache involved.


$ for n in r1 r2 r3 r4; do ssh $n time dd if=`pwd`/all_card_trans.ul of=/dev/null bs=1M &; done
[5] 13325
[6] 13326
[7] 13327
[8] 13328
$ 32458+1 records in
32458+1 records
34034910300 bytes (34 GB) copied, 154.117 seconds, 221 MB/s

real    2m34.127s
user    0m0.014s
sys     0m3.073s
32458+1 records in
32458+1 records out
34034910300 bytes (34 GB) copied, 155.113 seconds, 219 MB/s

real    2m35.123s
user    0m0.020s
sys     0m3.127s
32458+1 records in
32458+1 records out
34034910300 bytes (34 GB) copied, 155.813 seconds, 218 MB/s

real    2m35.821s
user    0m0.026s
sys     0m3.210s
32458+1 records in
32458+1 records out
34034910300 bytes (34 GB) copied, 155.89 seconds, 218 MB/s

real    2m35.901s
user    0m0.017s
sys     0m3.039s

With Exadata in mind, the idea is to offer a comprehensive solution for data warehousing. All too often I see data loading claims that start with the flat files sort of magically appearing ready to be loaded. Oh no, we don’t think that way. The data is outside on a provider system somewhere and has to be staged in advance of ETL/ELT. Since DBFS exploits the insane bandwidth of Exadata, it is an extremely good data staging solution. The data has to be rapidly ingested into the staging area and then rapidly loaded. A bottleneck on either part of that equation will be your weakest link.

Just think, no external systems required for data staging. No additional storage connectivity, administration, tuning, etc.

And, yes, it can do more than a single dd process on each node! Much more.

Exciting stuff.

Webcast Announcement Clarification. Exadata Technical Deep Dive Part IV.

I just noticed that my announcement for the up-coming Part IV in my Exadata Technical Deep Dive series was missing the date and time. So, here it is:

Thursday, June 18, 2009 12:00 PM – 1:00 PM CDT

Oracle Data Warehouse Performance Issues? Solve It The Old-Fashioned Way With A Third-Party Accelerator!

I read Curt Monash’s report on the current state of affairs at Dataupia and it got me thinking. I agree with Curt on his position toward add-on or external accelerator-type technology. See, one of Dataupia’s value propositions was to accelerate I/O for Oracle Database External Tables. To the best of my knowledge it basically offered high bandwidth, cached flat file access.

About this time last year I produced a bit of a toungue-in-cheek post about Dataupia. A blog reader posted the following comment on that thread:

… This product is very, very real.

It works as an external table in Oracle so it’s transparent to all your BI tools. They do a lot of work with SQL that Oracle passes to make it usable.

You have to re-point your ETL loads at Dataupia directly but they should run with very little alteration.

Speed is 10x Oracle at these volumes (2Tb+).

As most folks know I was deep into HP Oracle Exadata Storage server performance work at that time and couldn’t really go toe-to-toe with any of the DW/BI appliance or accelerator folks. Oracle had not yet released Exadata. The idea of accelerating Oracle ten-fold is certainly no longer all that avant-garde given the proven acceleration Oracle Exadata Storage Server provides.

What I wanted to point out at the time is that accelerating the loading of an Oracle Data Warehouse is indeed important, but surely not sufficiently critical to warrant bringing in another vendor and working out all the plumbing. I had a suspicion then that the blog reader who posted that comment was not fully aware that the value proposition supposedly went beyond accelerating ETL to offering run-time access to the flat files they housed in their Satori server. Yes, running queries against External Tables just because they offer a lot of cache and a lot of I/O bandwidth. At least that is what I got from reading their datasheet.

Erroneously Accelerating Accelerates What?
The problem with that story is that query throughput from External Tables is very seldom an I/O issue. See, scanning External Tables requires conversion from ASCII flat-file text to Oracle data types on the fly. To that end, scanning External Tables is a CPU-intensive task. For instance, if you load data from an External Table into a data warehouse (internal, true) table and then compare scan throughput of both you’ll see that processor saturation will impede the External Table scan. Same-query comparisons commonly show 80% lower throughput accessing External Tables compared to internal tables and I’m not talking about an I/O-hobbled External Table comparison. That, of course, depends on the processor bandwidth available to such a test because the less processor bandwidth available, the more significant the skew towards the internal table.  What I’m trying to say is that if you accelerate External Table I/O, say, 10x, you need as much as 10x more processor bandwidth to handle it. So, sure, if you take a totally I/O bound query and do something like this External Table acceleration technique, you will see significant performance increase. On the contrary, a host processor-bound situation will not benefit from this sort of accelerator. Architecture…it’s important.

Tacking on accelerators is just not a reasonable approach. I recall a lot of hoopla back in about 2005 or so about another one of these sorts of external accelerator offerings—Xprime. I don’t hear much about them any more, other than perhaps bits and pieces about intellectual property infringement claims against DATAllegro (Microsoft).

I’m no stranger to the external acceleration game, but I have generally steered clear of such approaches. I have always leaned toward a more native approach. Offer a better platform, not an external platform. About the same time Xprime was garnering quite a bit of interest, we at my former company, PolyServe, had been putting the final touches on product infrastructure that offered scale-out reporting using clustering technology. Of course the story had all the common tag words such as transparent, seamless, scalable, etc. Unlike usual, however, the claims were true. But, no matter. Nobody cared. I sure thought people would have clamored for up to 16-fold throughput increase for processor-intensive reporting jobs. Oh well…memories. It was an interesting project to work on though as this old paper I wrote suggests.

So What Does This Have To Do With Exadata?
It’s probably about high time people stop getting venture capital to “solve” a “problem” that Oracle Database supposedly has with data warehouse workloads.

Oracle Exadata Storage Server Technical Deep Dive – Part IV

BLOG UPDATE (18-JUN-2009): Links to the recorded webcasts can be found in my Papers, etc section. The original blog post follows:

We’ve set the date for Part IV. As an aside, I’m sorry to report that the recorded webcast of Part III is still not available from IOUG.

I recommend that folks view Part I as a minimum prerequisite for Part IV. I won’t spend any time in review. You can access Part I and Part II here.

Announcing Part IV:

Kevin Closson will continue his “Technical Deep Dive” series in Part IV by covering:

– Loading the Data Warehouse in an Exadata Environment
* The Data Staging Model
* A Data Loading Performance Study

Space is limited.
Reserve your Webinar seat now at:

https://www1.gotomeeting.com/register/473487440

World-Record TPC-H Results Require World-Record Floor Space?

This may be one of the quickest follow-ups to one of my own posts. I just saw an IBM blogger’s tongue-in-cheek post about the World-Record Oracle Database 11g TPC-H result. The post reads:

Yesterday, for the very first time, I went to Costco.

Now for those of you who live on Mars, Costco is one of those big warehouse type member only stores. You can get almost anything there but can never be sure exactly what will be there. I ended up with two of the biggest cans of tuna I have ever seen, a jar of Kalamata olives as big as a fishbowl, and enough toilet paper for a year.

Which reminded me of yesterday’s new TPC-H BI result from HP. HP now leads in the 1000GB space here – by using 64 servers with 512 cores. And 6 humongously specialized storage devices.(1)

It’s fun to think about the floor space, energy, and resources to manage that infrastructure. At least the toilet paper can go in the basement.

I got a chuckle out of that post and would have just commented on the blog, but there was some sort of login credential required to comment.

So What is My Comment?
Well, according to the blog header, the blogger who posted this humorous bit is Chief Technical Strategist, Performance Marketing for the IBM Systems and Technology Group. I think since a professional holding such a position as this seems to have missed a couple of critical points, I thought I’d point out a couple of things.

The blogger referred to the 6 HP Oracle Exadata Storage Server cells as “humongously specialized storage devices.” Yes, Exadata is humongously, enormously, gigantically, immensely, vastly, colossally specialized. However, the blogger moved on to insinuate there would be floor space issues with such a beast.

In case anyone else missed the point, this was 4 10U HP BladeSystem C7000 enclosures. That’s 40U. The humongously specialized, but minimally sized, storage devices were 6 2U HP Oracle Exadata Storage Servers. Sure, there were a couple of switches and some other such supporting gear, but, honestly, is it that “fun to think about” the floor space required for 52U worth of kit?  🙂


DISCLAIMER

I work for Amazon Web Services. The opinions I share in this blog are my own. I'm *not* communicating as a spokesperson for Amazon. In other words, I work at Amazon, but this is my own opinion.

Enter your email address to follow this blog and receive notifications of new posts by email.

Join 819 other subscribers
Oracle ACE Program Status

Click It

website metrics

Fond Memories

Copyright

All content is © Kevin Closson and "Kevin Closson's Blog: Platforms, Databases, and Storage", 2006-2015. Unauthorized use and/or duplication of this material without express and written permission from this blog’s author and/or owner is strictly prohibited. Excerpts and links may be used, provided that full and clear credit is given to Kevin Closson and Kevin Closson's Blog: Platforms, Databases, and Storage with appropriate and specific direction to the original content.