Archive for the 'oracle' Category



Why I Don’t Link To Official Oracle Corporate Blogs

I recently  realized I’ve failed to link to the official Oracle blog on data warehouse related topics. For whatever reason I thought I did this some time back but have apparently been speed-reading my blogroll. Better late than never…

I’m adding The Data Warehouse Insider to my blog roll.

You might find this post about In-database Map/Reduce interesting.

Why Doesn’t Oracle Offer Exadata Related Partner Programs?

…there is, in fact, a new Partner Program…

This is just a quick blog entry to point folks to what I think is a very important Oracle press release.

On October 12, Oracle announced the creation of  the Oracle Exadata Partner Program. While at Open World this week I spoke with a few of my friends that have very successful Oracle consultancies. As usual they asked me how they might get into an appropriate program whereby they can help customers chose when Exadata will fit their requirements. I think this program souns like a great start.

I captured the following list from the press release:

  • Following the announcement of Oracle® Exadata V2 last month, Oracle today announced the Oracle Exadata Partner Program.
  • This new program will allow Oracle partners to resell Sun Oracle Database Machines and Sun Oracle Exadata Storage Servers, and also provides partners with enablement resources to help build value added services for Oracle’s customers.
  • Resellers must be enrolled in the Oracle PartnerNetwork (OPN) and hold a valid Full Use Distribution Agreement to resell Oracle Exadata products.
  • Solution Providers and Systems Integrators are encouraged to build Oracle Exadata expertise and implementation services around Business Intelligence, Data Warehousing, VLDB and OLTP environments as well as vertical industry expertise in the retail, financial services, communications, healthcare and public sector segments.
  • ISVs are invited to review the industry leading performance, scalability and reliability of Oracle Exadata V2 as the platform for their own Oracle based solutions.
  • As part of the new OPN Specialized Program, Oracle is developing the Oracle Exadata Knowledge Zone with Guided Learning Paths and partner centric training materials that are scheduled to become available in the coming weeks.
  • An Oracle Exadata Specialization is also expected to be available shortly after OPN Specialized systems’ projected go-live date of December 1. (See accompanying announcement)
  • Oracle Exadata V2, developed by Sun and Oracle, is the world’s fastest database machine capable of both data warehousing and online transaction processing (OLTP) applications.
  • The Oracle Exadata Partner Program will roll out over the next several months. Details and information are available on the OPN portal.

Sun Oracle Database Machine Cache Hierarchies and Capacities – Part 0.I

BLOG CORRECTION:  Well, nobody is perfect. I need to point out that I must  be too Exadata-minded these days. Exadata Smart Scan returns uncompressed data when performing a Smart Scan of a table stored in Hybrid Columnar Compression form. However, it was short-sighted for me to state categorically that the cited 400 GB of DRAM available as cache in the Sun Oracle Database Machine can only be used for uncompressed data. It turns out that the model in mind by the company for this cache is to buffer data not returned by Smart Scan but instead returned in simple block form and cached in the block buffer pool of the SGA on each of the 8 database servers. So, I was both right and wrong. The Sun Oracle Database Machine is a feature-rich product and I was too Exadata-centric with the information in this post. I have been properly spanked and I’m as contrite as I can possibly be. The original post follows:

…goofy title I know…but, hey, the Roman numeral system had no zero. This is a pre-“Deep Dive Series” post if you will.

I see my colleague Jean-Pierre Dijcks has a blog entry covering the Sun Oracle Database Machine features. It is a good overview but I need to point out a minor correction.

The piece suggests there is 400 GB aggregate DRAM cache in the database grid. There is indeed 576 GB aggregate DRAM cache available between the 8 database hosts so there should be no problem dedicated 70% of that to caching by the new Oracle Database 11g Release 2 Parallel Query caching feature. That particular feature came into play in the recent 1 TB scale record TPC-H result that I blogged about here. The nit I pick with the post is how cites 400 GB raw DRAM capacity and up to 4 TB user data DRAM capacity (presuming the commonly achievable Hybrid Columnar Compression ratio of roughly 10 to 1). Unfortunately, I haven’t had the time to produce one of my typical “Deep Dive Series” webcasts covering such matters as data flow and plumbing (but I’ll get to it) in the Sun Oracle Database Machine. In the meantime I need to point out that data flows from the intelligent storage grind into the database grid in uncompressed form when Oracle Exadata Storage Server cells are scanning (Smart Scan)  Hybrid Columnar Compression tables. So, the DRAM cache capacity of the database grid is an aggregate 400 GB. *Note: Please see the blog correction above regarding how this DRAM cache is populated to achieve the advertised, effective 10TB capacity.

Now, having said that, the Exadata Smart FLASH Cache does indeed cache data in its Hybrid Columnar Compression form. So the effective cache capacity is 50 TB in that tier  presuming a 10 to 1 compression ratio. Data flies  off of FLASH at an aggregate rate of 50 GB/s. Thank heavens there are 16 Xeon 5500 (Nehalem) processor threads in each cell to uncompress the data and perform filtration and column compression.

By the way, the 50 GB/s is actually a safe, conservative number as it represents roughly 900 MB/s per FLASH card each of which has a dedicated 1GB/s lane into memory.

Announcement: Fresh Oracle White Paper Covering Sun Oracle Exadata Storage Server and Database Machine

I’ve been getting quite a few emails from folks asking why I haven’t been posting content about Sun Oracle Database Machine. One such reader asked:

You guys must not actually have any real proof of this stuff or else you would be blogging for sure

I’m up to my neck in Sun Oracle Database Machine testing and performance characterization. Keeping several of these high-end beasts busy under my intense scrutiny is a good piece of work. That’s exactly why I haven’t been sticking my head up out of the foxhole and blogging.

While I have a lot of material I could blog about I thought it would be proper for everyone to get the official story in the form of Oracle white papers and the many slides to be shown by executives and upper management at OpenWorld 2009 before I started in with my incoherent ramblings and trivial pursuit. One such white paper can be found at the following URL:

A Technical Overview of the Sun Oracle Exadata Storage Server and Database Machine

Some readers have been asking me to produce Sun Oracle Exadata Storage Server FAQ-style posts in the same manner I did for the HP Oracle Exadata Storage Server and HP Oracle Database Machine. I have a stack of those ready to go but need time to put them out.

New Addition To My Blogroll: James Morle.

OakTable Network co-founder and very old friend of mine, James Morle, has started a blog. I harbor a great deal of respect for James and am looking forward to his posts. You can find his blog at the following link:

jamesmorle.wordpress.com

I should also point out that James wrote a very, very good book way back in the Oracle8i days. While that may seem old and irrelevant, I wager 99.42% of all folks reading this blog can learn from that book—today! You can find the book here.

Announcement: Winter Corporation Report Covering Improved Oracle Database 11g Release 2 Real Application Clusters Manageability

Winter Corporation published a report covering Oracle Database 11g Release 2 improvements in Real Application Clusters manageability.

There Can’t Be Too Much Information Offered About Sun Oracle Database Machine At Open World 2009, Right?

BLOG UPDATE 21-SEP-2009: The session Glenn Fawcett and I were scheduled to deliver has been  cancelled.

They are letting me out of my cage long enough to attend Open World 2009. I’ll be working some of the Sun Oracle Database Machine demos and offering a couple of low-key sessions. One of the sessions is a joint-session with my old friend Glenn Fawcett. Glenn and I have been doing some performance engineering work on a full-rack Sun Oracle Database Machine. I don’t yet know the time slot for that session but I’ll post it here when I find out.

I’ve also signed up to deliver a session on Monday October 12 in the Open World UnConference. I signed up for it before the Sun Oracle Database Machine announcement so I gave the title of the session a bit of a stealth-title. I’ll be talking about Exadata, but perhaps more importantly I’ll have a lengthy question and answer session. If you check out the schedule you’ll see my session is in the same room following two more interesting sessions by my friend, co-worker and fellow OakTable Network member Greg Rahn and fellow OakTable Network member, and luminary, Cary Milsap:

1pm
Overlook II: Chalk & Talk: The Core Performance Fundamentals Of Oracle Data Warehousing (Greg Rahn, Database Performance Engineer, Real-World Performance Group @ Oracle)

2pm
Overlook I: Fundamentals of Performance (Oracle ACE Director Cary Millsap)

3pm
Overlook II: Oracle Exadata Storage Server FAQ Review and Q&A with Kevin Closson (Performance Architect, Oracle)

Oracle Drops Exadata In Favor of Sun FlashFire Based OLTP Database Machine?

I’ve had numerous emails questioning where Exadata plays in the yet to be announced OLTP Database Machine. The link I offered in my previous post does not in fact mention Exadata so I understand the emails.

The following is another link regarding the announcement where a depiction of Exadata is prominently featured. Folks, it is Exadata.

Announcing the World’s First OLTP Database Machine with Sun FlashFire Technology

One For Your Calendar

Larry Ellison to Announce Sun Oracle Database Machine Specialized for OLTP.

World-Record TPC-H Result Proves Oracle Exadata Storage Server is 10x Faster Than Conventional Storage on a Per-Disk Basis?

BLOG UPDATE (02-Feb-2010): This post has confused some readers. I make mention in this post how Exadata Storage Server does not cache data. Please remember that the topic of this post is an audited TPC-H result that used Version 1 Exadata Storage Server cells. Version 2 Exadata Storage Server is the first release that caches data (the read-only Exadata Smart Flash Cache).

I’d like to tackle a couple of the questions that have  come at me from blog readers about this benchmark:

Kevin, I saw the 1TB TPCH benchmark number. It is very huge. You say Exadata does not cache data so how can it get such result?

True, I do say Exadata does not cache data. It doesn’t. Well, there is a .5 GB write cache in each cell, but that doesn’t have anything to do with this benchmark. This was an in-memory Parallel Query benchmark result. The SGA was used to cache the tables and indexes. That doesn’t mean there was no physical I/O (e.g., sort spilling, etc), but the audited runs were not a proof-point for scanning tables or indexes with offload processing.

Under The Covers
There were 6 HP Oracle Exadata Storage Servers (cells) in the configuration. Regular readers therefore know that there is no more than 6 GB up-wind bandwidth regardless of whether or not the data is cached in the cells. The database grid in this benchmark had 512 Xeon 5400 processor cores. I assure you all that 6 GB/s cannot properly feed 512 of such processor cores since that is only 12 MBPS/core.

Let me just point out that this result is with Oracle Database 11g Release 2 on a 64-node database grid with an aggregate memory capacity of roughly 2TB. The email continued:

I guess this prove Oracle with Exadata is 10x faster?

I presume the reader was referring to the  result in the prior Oracle Database 11g 1TB TPC-H with conventional storage. Folks, Exadata can be 10x faster than Oracle on state of the art conventional storage (generally misconfigured, poorly provisioned, etc). No argument here. But, honestly, I can’t sit here and tell you that 6 Exadata cells with 72 disks is 10x faster than 768 15K RPM drives connected via 128 4Gb Fibre Channel ports used in the prior Oracle 1TB result since that is about 50 GB/s theoretical I/O bandwidth. If you investigate that prior Oracle Database 11g 1TB TPC-H result you’ll see that it was configured with less than 20% of the RAM used by the new Oracle Database 11g Release 2 result (2080 GB aggregate vs 384 GB).

So, what’s my point?

This new world-record is a testimonial to the scalability of Real Application Clusters for concurrent, warehouse-style queries. As much as I’d love to lay claim to the victory on behalf of Exadata, I have to point out, in fairness, that in spite of playing a role in this benchmark the result cannot be attributed to the I/O capability of Exadata.

In short, there is no magic in Exadata that makes 6 12-disk storage cells (72 drives) more I/O capable than 768 drives attached via 128 dual-port 4GFC HBAs.

I’m just comparing one Oracle Database 11g result to another Oracle Database 11g result to answer some blog readers’ questions.

So, no, Exadata is not 10x faster on a per-disk basis. Data comes off of round-brown spinning thingies at the same rate when downwind of Oracle via Exadata or Fibre Channel.  The common problem with conventional storage is the plumbing.  Balancing the producer-consumer relationship between storage and an Oracle Database grid with conventional storage even at the rate produced by a measly 6 Exadata Storage Server cells can be a difficult task. Consider, for example, that one would require a minimum of 15 active 4GFC host bus adapters to deal with 6GB/s. Grid plumbing requires redundancy so one would require and additional 15 4GFC paths through different ports and a different switch in order to architect around single points of failure. I’ve lived prior lives rife with FC SAN headaches and I can attest that working out 30 FC paths can be a real headache.

Using Linux /proc To Identify ORACLE_HOME and Instance Trace Directories.

I recently had a co-worker access one of my systems running Oracle Database 11g. He needed to poke around with focus on an area that he specialized in. After getting him access to the server I got an email from him asking where the trace files are for the instance he was investigating.

This is one of those Carry On Wayward Googler™  sort of posts. Most of you will know this, but it may help someone someday. It did help my co-worker as this was the way I answered his question.

You can find out a lot about an instance without even knowing which ORACLE_HOME it is executing out of by spelunking about in /proc. In the following text box you’ll see how to find the ORACLE_HOME and trace directories for an instance by looking at /proc/<PID>/fd and /proc/<PID>/exe of the LGWR process. This box had an instance called test and an ASM instance. So in this case the ORACLE_HOME values were /u01/app/oracle/product/11.2.0/dbhome_1 and /u01/app/11.2.0/grid.


$ ps -ef | grep lgwr | grep -v grep
oracle    3548     1  0 Sep02 ?        00:00:27 ora_lgwr_test3
oracle    8734     1  0 Sep02 ?        00:00:00 asm_lgwr_+ASM3
$
$ ls -l /proc/8734/exe /proc/3548/exe
lrwxrwxrwx 1 oracle oinstall 0 Sep  2 21:09 /proc/3548/exe -> /u01/app/oracle/product/11.2.0/dbhome_1/bin/oracle
lrwxrwxrwx 1 oracle oinstall 0 Sep  2 11:09 /proc/8734/exe -> /u01/app/11.2.0/grid/bin/oracle
$
$ ls -l /proc/8734/fd  /proc/3548/fd | grep trace | grep -v grep
l-wx------ 1 oracle oinstall 64 Sep  2 21:09 11 -> /u01/app/oracle/diag/rdbms/test/test3/trace/test3_ora_3501.trc
l-wx------ 1 oracle oinstall 64 Sep  2 21:09 12 -> /u01/app/oracle/diag/rdbms/test/test3/trace/test3_ora_3501.trm
l-wx------ 1 oracle oinstall 64 Sep  2 11:09 16 -> /u01/app/oracle/diag/asm/+asm/+ASM3/trace/+ASM3_ora_8636.trc
l-wx------ 1 oracle oinstall 64 Sep  2 11:09 17 -> /u01/app/oracle/diag/asm/+asm/+ASM3/trace/+ASM3_ora_8636.trm

Intel Xeon 5500 Nehalem: Is It 17 Percent Or 2.75-Fold Faster Than Xeon 5400 Harpertown? Well, Yes Of Course It Is!

I received two related emails while I was out recently for a couple of days of fishing and hiking. I thought they’d make for an interesting blog entry. The first email read:

…our tests show very little performance improvement on nehalem cpus compared to older Xeon…

And, the other email was the polar opposite:

…in most of our tests the Xeon 5500 was over 2 times as fast as the harpertown Xeon…

And the email continued:

…so we think you should stop saying that Xeon 5500 is double the perf of older xeon

Well, I can’t make everyone happy. I tend to say that Intel Xeon 5500 (Nehalem) processors are twice as fast as Harpertown Xeon (5400) as a conservative, well-rounded way to set expectations.

Introducing Fat and Skinny
OK, bear with me now, this is a wee tongue-in-cheek. The reader who emailed me with the report of near parity between Nehalem and Xeon is not lying, he’s just skinny. And the reader who admonished me for my usual low-ball citation of 2x performance vis a vis Nehalem versus Harpertown? No, he’s not lying either…he’s fat. Allow me to explain.

It’s really quite simple. If you run code that spends a significant portion of processor cycles operating on memory lines in the processor cache, you are operating code that has a very low CPI (cycles per instruction) cost. In my terminology such code is “skinny.” On the other hand code that jumps around in memory causing processor stalls for memory loads has a high CPI and is, in my terminology, fat.

Skinny code more or less relegates the comparison between Harpertown and Nehalem to one of clock frequency whereas fat code is really where the rubber hits the road. The more load and store hungry (fat) the code is the more the Nehalem pay-off will be.

Let’s take a look at two different, simple programs to help make the point. Using fat.c and skinny.c I’ll take timings on a Harpertown and Nehalem based boxes. As you can see, skinny.c simply hammers away on the same variable and does not leave L2 cache. On the other hand, fat.c treats its memory allocation as an array of 8-byte longs and skips to every 8th one in a loop in order to force memory loads since the cache line size on this box is 64 bytes. NOTE: do not compile these with -O (or change the longs in the array to volatile long). A simple gcc without args will suffice.

So, skinny.c has a very low CPI and fat.c has a very high CPI.

In the following examples, the model name field from cpuid output tells us what each system is. The E5430 is Harpertown Xeon and the 5570 is of course Nehalem. In terms of clock frequency, the Nehalem processors are 10% faster than the Harpertown Xeons.

In the following box you’ll see screen-scrapes I took from two different systems, one based on Nehalem and the other Harpertown. Notice how skinny only improves by 17% with the same executable on Nehalem compared to Harpertown.


# cat /proc/cpuinfo | grep 'model name'
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
# md5sum skinny
df86d9a278ea33b7da853d7a17afdd46  skinny

# time ./skinny

real    6m3.658s
user    6m3.567s
sys     0m0.001s
#

# cat /proc/cpuinfo | grep 'model name'
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
# md5sum skinny
df86d9a278ea33b7da853d7a17afdd46  skinny
# time ./skinny

real    5m1.941s
user    5m2.043s
sys     0m0.001s

In the next box you’ll see screen-scrapes from the same two systems where I ran the “fat” executable. Notice how the Harpertown Xeon took 2.75x longer to process the fat.


# cat /proc/cpuinfo | grep 'model name' | head -1
model name      : Intel(R) Xeon(R) CPU           E5430  @ 2.66GHz
# md5sum fat
b717640846839413c87aedd708e8ac0d  fat
# time ./fat

real    1m57.731s
user    1m57.659s
sys     0m0.045s

# cat /proc/cpuinfo | grep 'model name' | head -1
model name      : Intel(R) Xeon(R) CPU           X5570  @ 2.93GHz
# md5sum fat
b717640846839413c87aedd708e8ac0d  fat
# time ./fat

real    0m42.834s
user    0m42.803s
sys     0m0.023s

So, as it turns out, we can believe both of the folks that sent me email on the matter.

Oracle Switches To Columnar Store Technology With Oracle Database 11g Release 2?

I see fellow OakTable Network member Tanel Poder has blogged that Oracle Database 11g Release 2 has switched to offer columnar store technology. Or at least one could infer that from the post.

I left a comment on Tanel’s blog but would like to make a quick entry here on the topic as well. Oracle Database 11g Release 2 does not offer column-store technology as thought of in the Vertica (for example) sense. The technology, available only with Exadata storage, is called Hybrid Columnar Compression. The word hybrid is important.

Rows are still used. They are stored in an object called a Compression Unit. Compression Units can span multiple blocks. Like values are stored in the compression unit with metadata that maps back to the rows.

So, “hybrid” is the word. But, none of that matters as much as the effectiveness. This form of compression is extremely effective.

Less Blogging? A Mixed Blessing.

I have been down in the foxhole (lab work) non-stop and haven’t come up for air except to take a few short days off. Some of you may enjoy photos of one of the desert sunsets I took in after a full day of fishing… ah, sighs of relief!

It was a true pleasure watching this sunset go from ice to fire. It’s hard to beat desert sunsets…

IMG_3245

IMG_3252

IMG_3257

IMG_3261

Intel Xeon 5500 (Nehalem EP) NUMA Versus Interleaved Memory (aka SUMA): There Is No Difference! A Forced Confession.

I received an interesting email recently from a reader that takes offense at how I dare to discuss the differences between Intel Xeon 5500 (Nehalem) systems operating in NUMA versus SUMA/SUMO mode. One excerpt of the email read:

…and I think you are just creating confusion and chaos to gain popularity with your NUMA versus non-NUMA stuff. We tested everything we can think of and see no difference when booted with NUMA or non-NUMA…

I don’t doubt for one moment that the testing performed by this reader showed no performance differences between NUMA and SUMA because I have no idea whatsoever what his testing consisted of. And, besides, Xeon 5500 Nehalem EP is one extremely nice NUMA package. That is, when running non-NUMA aware software on this particular NUMA offering you can rest assured that you won’t likely fall over dead from NUMA pathologies. That’s good, but does that mean there really is no difference when booted in the NUMA versus SUMA? Hardly!

Please allow me to explain something. Intel Xeon 5500 (Nehalem) is a very tightly coupled NUMA system. Remote memory references are only about 20% more costly than local. If you measure a workload that does not saturate the processors you are very unlikely to detect any difference in throughput. If you have a program that only drives a processor core to, say, 80% utilization you will most likely not see any throughput difference if the process performs all its I/O into remote memory or local memory. When using only remote memory the process would consume moderately more processor cycles, however unless the code is overly-synthetic so as to force a high rate of L2 misses the result would likely be equivalent throughput in both the local and remote cases.

NUMA/SUMA: The Ever-Hypothetical Topic
Let’s stop talking in the hypothetical. How about something that, gasp, real Oracle Database Administrators have to do more than just occasionally.  Consider for a moment transferring a sizable zipped ASCII file in preparation for loading into an Oracle Data Warehouse. When booting in the default NUMA mode and running Linux, memory is presented to processes in multiple hierarchies. For example, the following box shows a freshly booted Intel Xeon 5500 (Nehalem EP) box with 16 GB total RAM segmented into two memories. Notice how just 7 minutes after booting up memory has been consumed in a non-symmetrical fashion. The numactl command shows that roughly 40% more memory has been allocated from node 0 memory compared to node 1. That’s because not every memory usage in the Linux kernel (including drivers) is NUMA aware. But that is not what I’m blogging about.

# uptime;numactl --hardware
 13:28:30 up 7 min,  1 user,  load average: 0.00, 0.09, 0.07
available: 2 nodes (0-1)
node 0 size: 8052 MB
node 0 free: 5773 MB
node 1 size: 8080 MB
node 1 free: 7955 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10
# cat /proc/meminfo
MemTotal:     16427752 kB
MemFree:      14059424 kB
Buffers:         19588 kB
Cached:         239480 kB
SwapCached:          0 kB
Active:          66308 kB
Inactive:       217152 kB
HighTotal:           0 kB
HighFree:            0 kB
LowTotal:     16427752 kB
LowFree:      14059424 kB
SwapTotal:     2097016 kB
SwapFree:      2097016 kB
Dirty:            1848 kB
Writeback:           0 kB
AnonPages:       24408 kB
Mapped:          15024 kB
Slab:           170920 kB
PageTables:       3512 kB
NFS_Unstable:        0 kB
Bounce:              0 kB
CommitLimit:  10310892 kB
Committed_AS:   382752 kB
VmallocTotal: 34359738367 kB
VmallocUsed:    381716 kB
VmallocChunk: 34359356623 kB
HugePages_Total:     0
HugePages_Free:      0
HugePages_Rsvd:      0
Hugepagesize:     2048 kB
# free
             total       used       free     shared    buffers     cached
Mem:      16427752    2370064   14057688          0      19872     239852
-/+ buffers/cache:    2110340   14317412
Swap:      2097016          0    2097016

In this section of this blog entry I’d like to show a practical example of honest-to-goodness, real world work that doesn’t exhibit totally benign NUMA characteristics. Within a VNC I opened two xterm sessions. I’ll call them “left” and “right.” In the left xterm I’ll list a zipped ASCII file to capture the inode so as to prove my testing is happening against the same file. The file is inode 1701506. You’ll also see a stupid little script called henny_penny.sh named appropriately as I apparently come off as Henny Penny to folks like the reader who emailed me. The henny_penny.sh script executed in the left xterm showed that a shell with a parent process id of 23283 was able to sling the contents of all_card_trans.ul.gz into /dev/null at the rate of 4.9 GB/s. That is very fast indeed. It is that fast, in fact, because the file has been moved into the current directory with FTP so the contents of the approximately 1.5 GB file is cached in memory. Ah, but the question is, what memory?

# ls -li all* henny_penny*
1701506 -rw-r--r-- 1 root root 1472114768 Aug 14 11:31 all_card_trans.ul.gz
1701513 -rwxr-xr-x 1 root root         90 Aug 14 12:17 henny_penny.sh
# cat henny_penny.sh
ps -f
ls -li all_card_trans.ul.gz
date
dd if=all_card_trans.ul.gz of=/dev/null bs=1M
date
# sh ./henny_penny.sh
UID        PID  PPID  C STIME TTY          TIME CMD
root     23283 23280  0 12:13 pts/0    00:00:00 -bash
root     23849 23283  0 12:18 pts/0    00:00:00 sh ./henny_penny.sh
root     23850 23849  0 12:18 pts/0    00:00:00 ps -f
1701506 -rw-r--r-- 1 root root 1472114768 Aug 14 11:31 all_card_trans.ul.gz
Fri Aug 14 12:18:12 PDT 2009
1403+1 records in
1403+1 records out
1472114768 bytes (1.5 GB) copied, 0.30021 seconds, 4.9 GB/s
Fri Aug 14 12:18:12 PDT 2009

In the following box you’ll see how things behaved in the right xterm. I invoked henny_penny.sh (parent PID 23422) and voila dd(1) was able to shovel the contents of all_card_trans.ul.gz into /dev/null at a rate of 6 GB/s. Now, that’s only 22% faster for a totally memory-bound, CPU-saturated task so why would anyone other than Henny Penny care? Notice how the henny_penny.sh script included the output of the date(1) command. Just three seconds after “left” was muddling through at 4.9 GB/s, “right” proceeded to  rip through at 6.0 GB/s. Yes, memory hierarchy matters.

# sh ./henny_penny.sh
UID        PID  PPID  C STIME TTY          TIME CMD
root     23422 23420  0 12:14 pts/3    00:00:00 -bash
root     23856 23422  0 12:18 pts/3    00:00:00 sh ./henny_penny.sh
root     23857 23856  0 12:18 pts/3    00:00:00 ps -f
1701506 -rw-r--r-- 1 root root 1472114768 Aug 14 11:31 all_card_trans.ul.gz
Fri Aug 14 12:18:15 PDT 2009
1403+1 records in
1403+1 records out
1472114768 bytes (1.5 GB) copied, 0.244703 seconds, 6.0 GB/s
Fri Aug 14 12:18:15 PDT 2009

How, What, Why?
The left xterm and its children happen to be executing on cores 0-3 (SMT disabled at the moment but no matter) and the right xterm on cores 4-7. The FTP process executed on one (or more) of cores 4-7 and since Linux prefers to allocate buffers to a process such as this from local memory, you can see why henny_penny.sh in the right xterm achieved the throughput it did.

Who Cares?

Likely nobody until the Xeon 5500 Linux production uptake actually starts! In the meantime there is me (Henny Penny) and a few curiously morbid (er, uh, morbidly curious) Googlers who might stumble upon this trivia.

What’s This Have To Do With Nehalem EX?
Well, even the 4-socket Nehalem EX packaging implements single-hop remote memory. That’s a significant difference from the way 4-sockets were done with HyperTransport. So, I actually don’t expect NUMAisms such as this to be any more painful than with EP (2 socket).

I Still Think He’s Henny Penny
So, let’s take another look at this topic. I’ve already mentioned that Linux likes to allocate memory close to processes when running on Nehalem systems. That’s good, isn’t it? Well, the answer is yes, of course, it depends.

In the following text box you’ll see how I depleted free memory (down to 40MB free) from node 0 by writing zeros to a file. Consider yet another hypothetical with me for one moment. What happens when I execute, say, 100 processes that each allocates a moderate 16 MB of memory with malloc(3)?  Do you think Linux will yank these processes from me, their parent, and place them on node 1 or will they be homed on node 0 with their heaps allocated from node 1? Will it matter? What if they are producers and I am their consumer? Where should they execute? What if they each work on 1/100th of the dumb_test.out file reading into their respective heap? Well, at this point there is no way for 100 processes on node 0 (socket 0) to attack 1/100th segments (buffering in their heap) of that file without 100% remote memory overhead. Could such a “bizarre” hypothetical happen in production? Sure. Is there any way to properly deal with such an issue? Well, yes and no.

If the hypothetical “1/100th program” was coded to libnuma then it can assure process placement and therefore local heap. However, what about the fact that my work file is buffered entirely on node 0 memory? Wouldn’t that guarantee 100% local access to node 0 users of that file but 100% remote for node 1 users? Yes. That’s great for the node 0 users you might say. However, those node 0 users had better not malloc(3) any memory because you know where that memory is going to come from. ‘Round and ’round we go…

# numactl --hardware
available: 2 nodes (0-1)
node 0 size: 8052 MB
node 0 free: 5946 MB
node 1 size: 8080 MB
node 1 free: 7987 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10
# time dd if=/dev/zero of=dumb_test.out bs=1M count=5946;numactl --hardware
5946+0 records in
5946+0 records out
6234832896 bytes (6.2 GB) copied, 6.07315 seconds, 1.0 GB/s

real    0m6.091s
user    0m0.003s
sys     0m6.069s
available: 2 nodes (0-1)
node 0 size: 8052 MB
node 0 free: 40 MB
node 1 size: 8080 MB
node 1 free: 7652 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10

So, what if I cloak my test with libnuma attributes (inherited by dd from numactl(8))?  In the following text box you’ll see that instead of a Cyclops, memory was allocated nice and evenly from the page cache when I filled out the dumb.test.out file. So in this model, processes homed on either node 0 or node 1 are guaranteed a 50% local access rate when accessing dumb_test.out and I am protected from memory imbalances. In fact, if it was my system and had to stay with NUMA, I’d consider invoking shells under numactl –interleave. As such any non-NUMA aware programs (like FTP) will be granted memory in a round-robin fashion but any NUMA aware program (coded to libnuma calls) will execute as it would without being wrapped with numactl. It’s just a thought. It isn’t any official recommendation and, as my email in-box suggests, it doesn’t matter anyway…nonetheless, I think the following looks better than a cyclops:

# numactl --interleave=0,1 /bin/bash
# numactl -s
policy: interleave
preferred node: 0 (interleave next)
interleavemask: 0 1
interleavenode: 0
physcpubind: 0 1 2 3 4 5 6 7
cpubind: 0 1
nodebind: 0 1
membind: 0 1
# numactl --hardware
available: 2 nodes (0-1)
node 0 size: 8052 MB
node 0 free: 5957 MB
node 1 size: 8080 MB
node 1 free: 7988 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10
# dd if=/dev/zero of=dumb_test.out bs=1M count=5957
5957+0 records in
5957+0 records out
6246367232 bytes (6.2 GB) copied, 6.24962 seconds, 999 MB/s
# numactl --hardware
available: 2 nodes (0-1)
node 0 size: 8052 MB
node 0 free: 2825 MB
node 1 size: 8080 MB
node 1 free: 4854 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10

DISCLAIMER

I work for Amazon Web Services. The opinions I share in this blog are my own. I'm *not* communicating as a spokesperson for Amazon. In other words, I work at Amazon, but this is my own opinion.

Enter your email address to follow this blog and receive notifications of new posts by email.

Join 819 other subscribers
Oracle ACE Program Status

Click It

website metrics

Fond Memories

Copyright

All content is © Kevin Closson and "Kevin Closson's Blog: Platforms, Databases, and Storage", 2006-2015. Unauthorized use and/or duplication of this material without express and written permission from this blog’s author and/or owner is strictly prohibited. Excerpts and links may be used, provided that full and clear credit is given to Kevin Closson and Kevin Closson's Blog: Platforms, Databases, and Storage with appropriate and specific direction to the original content.