From pjones at whamcloud.com Tue Feb 11 22:05:07 2020
From: pjones at whamcloud.com (Peter Jones)
Date: Tue, 11 Feb 2020 22:05:07 +0000
Subject: [lustre-devel] Lustre 2.12.4 released
Message-ID: <5BC865C0-22D4-44A1-BBA8-A3EC9BCB0115@ddn.com>
We are pleased to announce that the Lustre 2.12.4 Release has been declared GA and is available for download. You can also grab the source from git .
Details of changes since 2.12.3 can be found in the 2.12.4 change log.
There are the following notable enhancements over 2.12.3:
RHEL 8.1 is now supported for clients (LU-12637) . Note that a kernel issue with cgroups has been found in our testing with RHEL 8.1 Lustre clients (LU-13063). We have been advised that this issue will be fixed in RHEL 8.2.
Please log any issues found in the issue tracking system.
Thanks to all those who have contributed to the creation of this release.
The next LTS release will be 2.12.5
-------------- next part --------------
An HTML attachment was scrubbed...
URL:
From green at whamcloud.com Wed Feb 12 06:22:23 2020
From: green at whamcloud.com (Oleg Drokin)
Date: Wed, 12 Feb 2020 06:22:23 +0000
Subject: [lustre-devel] New tag 2.13.52
Message-ID:
Hello!
I just tagged 2.13.52 in Lustre development branch. Here’s the changelog:
Alex Zhuravlev (5):
LU-13098 ptlrpc: supress connection restored message
LU-13130 tests: sanity-scrub to use full device size with ZFS
LU-12133 osd-zfs: set blocksize to 8K for llog objects
LU-12988 ldiskfs: skip non-loaded groups at cr=0/1
LU-12988 ldiskfs: mballoc to prefetch groups
Alexander Boyko (2):
LU-13093 osd: fix osd_attr_set race
LU-12593 osd: up i_append_sem during errors
Alexander Zarochentsev (1):
LU-13128 osc: glimpse and lock cancel race
Alexey Lyashkov (3):
LU-12214 selinux: Remove concatenating of selinux context
LU-12991 lnet: lnet response entries leak
LU-13036 lnet: avoid extra memory consumption
Amir Shehata (1):
LU-13049 lnet: peer lookup handle shutdown
Andreas Dilger (11):
LU-12865 tests: fix sanity 160f to be more robust
LU-8066 lfsck: use underscores in lfsck status files
LU-12470 tests: increase pdirops timeout
LU-12521 llapi: add separate fsname and instance API
Revert "LU-13120 build: Fix ZFS dependancies for osd-zfs-mount"
LU-11644 ptlrpc: show target name in req_history
LU-12518 llite: proper names/types for offset/pages
LU-12871 mdd: enable Changelog garbage collection
LU-13164 uapi: remove unused LUSTRE_DIRECTIO_FL
LU-13063 tests: remove checks for old RHEL versions
LU-13145 lnet: use conservative health timeouts
Andriy Skulysh (3):
LU-7791 ldlm: signal vs CP callback race
LU-13101 llite: eviction during ll_open_cleanup()
LU-13165 mdt: MSG_RESENT can be improperly cleared.
Arshad Hussain (1):
LU-12923 libcfs: Remove CLASSERT() for libcfs_private.h
Chris Horn (7):
LU-12756 lnet: Avoid extra lnet_remotenet lookup
LU-12756 lnet: Remove unused vars in lnet_find_route_locked
LU-12756 lnet: Refactor lnet_compare_routes
LU-12919 lnet: Fix source specified route selection
LU-13147 tests: Cleanup sanity-lnet on test failure
Revert "LU-12222 lnet: Check if we're sending to ourselves"
LU-12889 lnet: Do not assume peers are MR capable
Emoly Liu (1):
LU-12852 pfl: restrict the stripe count correctly
James Nunez (7):
LU-12928 tests: start running recovery-small 136
LU-13053 tests: fix conf-sanity call to umount_ldiskfs
LU-13063 tests: stop running sanity test 411
LU-1538 tests: standardize test script init – failover
LU-13194 tests: check server version sanityn 104
LU-10447 tests: deprecate use of $SETSTRIPE/$GETSTRIPE
LU-11607 tests: replace version/fstype calls in sanity/n
James Simmons (5):
LU-13119 osd-ldiskfs: set f_cred for app armour
LU-12822 uapi: properly pack data structures
LU-12977 ldiskfs: properly take inode_lock() for truncates
LU-9859 libcfs: move files out of libcfs/linux
LU-12598 osd-ldiskfs: always return errors for osd_ios_lf_fill
Jian Yu (1):
LU-12791 kernel: kernel update RHEL 8.0 [4.18.0-80.11.2.el8_0]
Jinshan Xiong (2):
LU-4198 clio: turn on lockless for some kind of IO
LU-4198 clio: AIO support for direct IO
Lai Siyao (3):
LU-13121 llite: fix deadlock in ll_update_lsm_md()
LU-13163 mdc: new kernel function xa_is_value()
LU-13191 osp: handle -EROFS in osp_sync_interpret()
Mikhail Pershin (4):
LU-13115 mdt: handle mdt_pack_sectx_in_reply() errors
LU-10664 tests: fix MPI tests in dom-performance.sh
LU-13136 dom: check read-on-open buffer presents in reply
LU-10198 llog: keep llog handle alive until last reference
Mr NeilBrown (25):
LU-9679 llite: fix possible race with module unload.
LU-13004 ptlrpc: Allow BULK_BUF_KIOV to accept a kvec
LU-13005 lnet: discard LNetEQGet and LNetEQWait
LU-10467 ptlrpc: refactor waiting in ptlrpc_set_wait()
LU-12678 socklnd: initialize the_ksocklnd at compile-time.
LU-12678 lnet: make "struct lnet_lnd" always "const".
LU-12678 lnet: remove locking protection ln_testprotocompat
LU-10467 obdclass: convert waiting in cl_sync_io_wait().
LU-9679 modules: use list_move were appropriate.
LU-13004 target: convert tgt_send_buffer to use KIOV
LU-12678 lnet: remove dead code: lnet_fini_locks()
LU-12678 lnet: fix small race in unloading klnd modules.
LU-12678 lnet: me: discard struct lnet_handle_me
LU-12678 socklnd: convert peers hash table to hashtable.h
LU-10467 lustre: convert most users of LWI_TIMEOUT_INTERVAL()
LU-10467 lustre: convert users of back_to_sleep()
LU-10467 ptlrpc: convert waiters on set->set_waitq
LU-10467 ldlm: convert waiting in ldlm_flock_completion_ast()
LU-10467 ldlm: convert waiting in ldlm_completion_ast()
LU-10467 ptlrpc: convert use of l_wait_event_exclusive_head()
LU-9679 general: add missing spaces to folded strings.
LU-9679 lnet: discard lnet_print_text_bufs()
LU-9679 lnet: use LIST_HEAD() for local lists.
LU-9679 lustre: use LIST_HEAD() for local lists.
LU-11300 lnet: remove lnd_query interface.
NeilBrown (6):
LU-8130 lu_object: factor out extra per-bucket data
LU-12460 llite: replace lli_trunc_sem
LU-12542 handle: remove locking from class_handle2object()
LU-12542 handle: use hlist for hash lists.
LU-12542 handle: discard h_lock.
LU-8304 libcfs: convert debug_ctlwq to a completion.
Olaf Faaland (1):
LU-11114 llite: Update mdc and lite stats on open|creat
Oleg Drokin (1):
New tag 2.13.52
Patrick Farrell (2):
LU-12518 llite: Accept EBUSY for page unaligned read
LU-11939 tgt: Do not assert during grant cleanup
Quentin Bouget (1):
LU-12806 llapi: use name_to_handle_at in llapi_fd2fid
Sebastien Buisson (2):
LU-13152 llapi: llapi_layout_get_by_xattr groks DoM
LU-13142 lod: cleanup layout checking
Serguei Smirnov (1):
LU-11385 odbclass: Handle gracefully if nsproxy is NULL
Shaun Tancheff (8):
LU-13039 quota: Ensure local buffer is null terminated
LU-13141 ldiskfs: block alloc performance patch
LU-12904 ldiskfs: Add ldiskfs support for linux 5.4
LU-12968 mgs: Prevent reading past end of buffer
LU-13120 build: Fix ZFS dependancies for osd-zfs-mount
LU-12861 libcfs: Cleanup use of bare printk
LU-12634 gss: uid_keyring and session_keyring moved
LU-13183 ldiskfs: Drop remove truncate warning patch
Swapnil Pimpale (2):
LU-3606 fsx: Add fallocate operation to fsx
LU-3606 lustre: Reserve OST_FALLOCATE(fallocate) opcode
Tatsushi Takamura (1):
LU-12287 lnet: handling device failure by IB event handler
Vitaly Fertman (1):
LU-11276 ldlm: fix lock convert races
Vladimir Saveliev (1):
LU-13099 lmv: disable statahead for remote objects
Wang Shilong (5):
LU-13092 lbuild: include lbuild-{fc,rhel,sles} to SIGNATURE
LU-13117 libcfs: fix to match right key in cfs_get_environ()
LU-13154 test: skip sanity-quota 66 if MDS version < 2.12.4
LU-13134 obdclass: use slab allocation for cl_dio_aio
LU-13180 lustre: reserve bit for RDMA-only memory RPC
From degremoa at amazon.com Fri Feb 14 17:13:48 2020
From: degremoa at amazon.com (Degremont, Aurelien)
Date: Fri, 14 Feb 2020 17:13:48 +0000
Subject: [lustre-devel] Setting GFP_FS flag for Lustre threads doing DMU
calls?
Message-ID: <15EB8039-20D2-447C-A9C8-8DCB455B862B@amazon.com>
Hello
I would like to bring a technical discussion undergoing for a ZFS patch which relates to Lustre.
Debugging a deadlock on an OSS we noticed a Lustre thread deadlocked itself due to memory reclaim, an arc_read() can trigger kernel memory allocation that in turn leads to a memory reclaim callback and a deadlock within a single zfs process. (see below for the full stack)
ZFS code should call spl_fstrans_mark() everywhere it could be doing memory allocation that could trigger ZFS cache reclaim. Doing so ended up adding GFP_FS flag for memory allocations in ZFS code.
After discussing this with Brian on https://github.com/zfsonlinux/zfs/pull/9987, there is a discussion wondering where is the good spot to add this. For proper layering, it seems they should rather be done in Lustre threads calling DMU calls, likewise this is done in ZPL for ZFS.
Brian said: "This will resolve the deadlock but it also somewhat violates the existing layering. Normally we call spl_fstrans_check() when setting up a new kthread if it's going to call the DMU interfaces, or for system calls it's done in our registered VFS callbacks. Feel free to update the PR, but before moving forward with this solution let's check with @adilger about potentially calling this on the Lustre side when they setup the threads which access the DMU. There may be other cases this doesn't cover."
What do you think of it?
PID: 108591 TASK: ffff888ee68ccb80 CPU: 12 COMMAND: "ldlm_bl_16"
#0 [ffffc9002b98adc8] __schedule at ffffffff81610f2e
#1 [ffffc9002b98ae68] schedule at ffffffff81611558
#2 [ffffc9002b98ae70] schedule_preempt_disabled at ffffffff8161184a
#3 [ffffc9002b98ae78] __mutex_lock at ffffffff816131e8
#4 [ffffc9002b98af18] arc_buf_destroy at ffffffffa0bf37d7 [zfs]
#5 [ffffc9002b98af48] dbuf_destroy at ffffffffa0bfa6fe [zfs]
#6 [ffffc9002b98af88] dbuf_evict_one at ffffffffa0bfaa96 [zfs]
#7 [ffffc9002b98afa0] dbuf_rele_and_unlock at ffffffffa0bfa561 [zfs]
#8 [ffffc9002b98b050] dbuf_rele_and_unlock at ffffffffa0bfa32b [zfs]
#9 [ffffc9002b98b100] osd_object_delete at ffffffffa0b64ecc [osd_zfs]
#10 [ffffc9002b98b118] lu_object_free at ffffffffa06d6a74 [obdclass]
#11 [ffffc9002b98b178] lu_site_purge_objects at ffffffffa06d7fc1 [obdclass]
#12 [ffffc9002b98b220] lu_cache_shrink_scan at ffffffffa06d81b8 [obdclass]
#13 [ffffc9002b98b278] shrink_slab at ffffffff811ca9d8
#14 [ffffc9002b98b338] shrink_node at ffffffff811cfd94
#15 [ffffc9002b98b3b8] do_try_to_free_pages at ffffffff811cfe63
#16 [ffffc9002b98b408] try_to_free_pages at ffffffff811d01c4
#17 [ffffc9002b98b488] __alloc_pages_slowpath at ffffffff811be7f2
#18 [ffffc9002b98b580] __alloc_pages_nodemask at ffffffff811bf3ed
#19 [ffffc9002b98b5e0] new_slab at ffffffff81226304
#20 [ffffc9002b98b638] ___slab_alloc at ffffffff812272ab
#21 [ffffc9002b98b6f8] __slab_alloc at ffffffff8122740c
#22 [ffffc9002b98b708] kmem_cache_alloc at ffffffff81227578
#23 [ffffc9002b98b740] spl_kmem_cache_alloc at ffffffffa048a1fd [spl]
#24 [ffffc9002b98b780] arc_buf_alloc_impl at ffffffffa0befba2 [zfs]
#25 [ffffc9002b98b7b0] arc_read at ffffffffa0bf0924 [zfs]
#26 [ffffc9002b98b858] dbuf_read at ffffffffa0bf9083 [zfs]
#27 [ffffc9002b98b900] dmu_buf_hold_by_dnode at ffffffffa0c04869 [zfs]
#28 [ffffc9002b98b930] zap_get_leaf_byblk at ffffffffa0c71e86 [zfs]
#29 [ffffc9002b98b988] zap_deref_leaf at ffffffffa0c720b6 [zfs]
#30 [ffffc9002b98b9c0] fzap_lookup at ffffffffa0c730ca [zfs]
#31 [ffffc9002b98ba38] zap_lookup_impl at ffffffffa0c77418 [zfs]
#32 [ffffc9002b98ba78] zap_lookup_norm at ffffffffa0c77c89 [zfs]
#33 [ffffc9002b98bae0] zap_lookup at ffffffffa0c77ce2 [zfs]
#34 [ffffc9002b98bb08] osd_fid_lookup at ffffffffa0b6f4ef [osd_zfs]
#35 [ffffc9002b98bb50] osd_object_init at ffffffffa0b68abf [osd_zfs]
#36 [ffffc9002b98bbb0] lu_object_alloc at ffffffffa06d9778 [obdclass]
#37 [ffffc9002b98bc08] lu_object_find_at at ffffffffa06d9b5a [obdclass]
#38 [ffffc9002b98bc68] ofd_object_find at ffffffffa0f860a0 [ofd]
#39 [ffffc9002b98bc88] ofd_lvbo_update at ffffffffa0f94cba [ofd]
#40 [ffffc9002b98bd40] ldlm_cancel_lock_for_export at ffffffffa0923ba1 [ptlrpc]
#41 [ffffc9002b98bd78] ldlm_cancel_locks_for_export_cb at ffffffffa0923e85 [ptlrpc]
#42 [ffffc9002b98bd98] cfs_hash_for_each_relax at ffffffffa05a85a5 [libcfs]
#43 [ffffc9002b98be18] cfs_hash_for_each_empty at ffffffffa05ab948 [libcfs]
#44 [ffffc9002b98be58] ldlm_export_cancel_locks at ffffffffa092410f [ptlrpc]
#45 [ffffc9002b98be80] ldlm_bl_thread_main at ffffffffa094d147 [ptlrpc]
#46 [ffffc9002b98bf10] kthread at ffffffff810a921a
Aurélien
From jsimmons at infradead.org Thu Feb 27 21:07:48 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:48 -0500
Subject: [lustre-devel] [PATCH 000/622] lustre: sync closely to 2.13.52
Message-ID: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
These patches need to be applied to the lustre-backport branch
starting at commit a436653f641e4b3e2841f38113620535e918dd3f.
Combining the work of Neil and myself this brings the lustre
linux client up to just btefore the he landing of Direct I/O
(LU-4198) support. Testing shows this work is pretty stable.
Alex Zhuravlev (19):
lustre: ptlrpc: idle connections can disconnect
lustre: osc: serialize access to idle_timeout vs cleanup
lustre: protocol: MDT as a statfs proxy
lustre: ptlrpc: new request vs disconnect race
lustre: ldlm: pass preallocated env to methods
lustre: mdc: use old statfs format
lustre: osc: re-check target versus available grant
lustre: ptlrpc: reset generation for old requests
lustre: osc: propagate grant shrink interval immediately
lustre: osc: grant shrink shouldn't account skipped OSC
lnet: libcfs: poll fail_loc in cfs_fail_timeout_set()
lustre: obdclass: put all service's env on the list
lustre: obdclass: use RCU to release lu_env_item
lustre: obd: add rmfid support
lustre: mdc: polling mode for changelog reader
lustre: llite: forget cached ACLs properly
lustre: ptlrpc: return proper error code
lustre: llite: statfs to use NODELAY with MDS
lustre: ptlrpc: suppress connection restored message
Alexander Boyko (9):
lustre: ldlm: fix l_last_activity usage
lustre: ptlrpc: don't zero request handle
lustre: mgc: don't proccess cld during stopping
lustre: llog: add startcat for wrapped catalog
lustre: llog: add synchronization for the last record
lustre: mdc: don't use ACL at setattr
lnet: adds checking msg len
lustre: llite: prevent mulitple group locks
lustre: obdclass: don't skip records for wrapped catalog
Alexander Zarochentsev (4):
lustre: llite: ll_fault should fail for insane file offsets
lustre: osc: don't re-enable grant shrink on reconnect
lustre: ptlrpc: grammar fix.
lustre: osc: glimpse and lock cancel race
Alexey Lyashkov (10):
lustre: lu_object: improve debug message for lu_object_put()
lnet: use right rtr address
lnet: use right address for routing message
lustre: mdc: reset lmm->lmm_stripe_offset in mdc_save_lovea
lustre: obdecho: reuse an cl env cache for obdecho survey
lustre: obdecho: avoid panic with partially object init
lustre: mgc: config lock leak
lnet: fix rspt counter
lnet: lnet response entries leak
lnet: avoid extra memory consumption
Alexey Zhuravlev (1):
lustre: grant: prevent overflow of o_undirty
Amir Shehata (87):
lnet: ko2iblnd: determine gaps correctly
lnet: refactor lnet_select_pathway()
lnet: add health value per ni
lnet: add lnet_health_sensitivity
lnet: add monitor thread
lnet: handle local ni failure
lnet: handle o2iblnd tx failure
lnet: handle socklnd tx failure
lnet: handle remote errors in LNet
lnet: add retry count
lnet: calculate the lnd timeout
lnet: sysfs functions for module params
lnet: timeout delayed REPLYs and ACKs
lnet: remove duplicate timeout mechanism
lnet: handle fatal device error
lnet: reset health value
lnet: add health statistics
lnet: Add ioctl to get health stats
lnet: remove obsolete health functions
lnet: set health value from user space
lnet: add global health statistics
lnet: print recovery queues content
lnet: health error simulation
lnet: lnd: conditionally set health status
lnet: router handling
lnet: update logging
lnet: lnd: Clean up logging
lnet: unlink md if fail to send recovery
lnet: set the health status correctly
lnet: Decrement health on timeout
lnet: properly error check sensitivity
lnet: configure recovery interval
lnet: separate ni state from recovery
lnet: handle multi-md usage
lnet: socklnd: improve scheduling algorithm
lnet: lnd: increase CQ entries
lnet: lnd: bring back concurrent_sends
lnet: use number of wrs to calculate CQEs
lnet: recovery event handling broken
lnet: clean mt_eqh properly
lnet: handle remote health error
lnet: setup health timeout defaults
lnet: fix cpt locking
lnet: detach response tracker
lnet: invalidate recovery ping mdh
lnet: fix list corruption
lnet: correct discovery LNetEQFree()
lnet: verify msg is commited for send/recv
lnet: select LO interface for sending
lnet: remove route add restriction
lnet: Discover routers on first use
lnet: use peer for gateway
lnet: lnet_add/del_route()
lnet: Do not allow deleting of router nis
lnet: router sensitivity
lnet: cache ni status
lnet: Cache the routing feature
lnet: peer aliveness
lnet: router aliveness
lnet: simplify lnet_handle_local_failure()
lnet: Cleanup rcd
lnet: modify lnd notification mechanism
lnet: use discovery for routing
lnet: MR aware gateway selection
lnet: consider alive_router_check_interval
lnet: allow deleting router primary_nid
lnet: transfer routers
lnet: handle health for incoming messages
lnet: misleading discovery seqno.
lnet: drop all rule
lnet: handle discovery off
lnet: handle router health off
lnet: push router interface updates
lnet: net aliveness
lnet: discover each gateway Net
lnet: look up MR peers routes
lnet: check peer timeout on a router
lnet: prevent loop in LNetPrimaryNID()
lnet: fix peer ref counting
lnet: honor discovery setting
lnet: warn if discovery is off
lnet: handle unlink before send completes
lnet: handle recursion in resend
lnet: discovery off route state update
lnet: o2iblnd: cache max_qp_wr
lnet: fix peer_ni selection
lnet: peer lookup handle shutdown
Andreas Dilger (55):
lustre: llite: increase whole-file readahead to RPC size
lustre: mdc: fix possible NULL pointer dereference
lustre: obdclass: allow specifying complex jobids
lustre: idl: remove obsolete directory split flags
lustre: obdecho: use vmalloc for lnb
lustre: mgc: remove obsolete IR swabbing workaround
lustre: mds: remove obsolete MDS_VTX_BYPASS flag
lustre: ptlrpc: fix return type of boolean functions
lustre: ptlrpc: remove obsolete OBD RPC opcodes
lustre: ptlrpc: assign specific values to MGS opcodes
lustre: ptlrpc: remove obsolete LLOG_ORIGIN_* RPCs
lustre: obdclass: remove unused ll_import_cachep
lustre: ptlrpc: add debugging for idle connections
lustre: mdc: move RPC semaphore code to lustre/osp
lustre: misc: name open file handles as such
lustre: osc: move obdo_cache to OSC code
lustre: idl: remove obsolete RPC flags
lustre: osc: clarify short_io_bytes is maximum value
lustre: misc: quiet console messages at startup
lustre: idl: use proper ATTR/MDS_ATTR/MDS_OPEN flags
lustre: lov: add debugging info for statfs
lustre: hsm: make changelog flag argument an enum
lustre: uapi: fix warnings when lustre_user.h included
lustre: ptlrpc: clean up rq_interpret_reply callbacks
lustre: lov: quiet lov_dump_lmm_ console messages
lustre: llite: remove cl_file_inode_init() LASSERT
lnet: libcfs: allow file/func/line passed to CDEBUG()
lustre: llite: enable flock mount option by default
lustre: lmv: avoid gratuitous 64-bit modulus
lustre: Ensure crc-t10pi is enabled.
lustre: lov: avoid signed vs. unsigned comparison
lustre: llite: limit statfs ffree if less than OST ffree
lustre: misc: delete OBD_IOC_PING_TARGET ioctl
lustre: misc: remove LIBCFS_IOC_DEBUG_MASK ioctl
lustre: obdclass: improve llog config record message
lustre: ptlrpc: allow stopping threads above threads_max
lustre: llite: improve max_readahead console messages
lustre: uapi: fix file heat support
lustre: mdt: improve IBITS lock definitions
lustre: obdclass: don't send multiple statfs RPCs
lustre: uapi: add unused enum obd_statfs_state
lustre: mdc: hold lock while walking changelog dev list
lustre: ptlrpc: make DEBUG_REQ messages consistent
lustre: obdclass: align to T10 sector size when generating guard
lustre: ptlrpc: fix watchdog ratelimit logic
lustre: llite: clear flock when using localflock
lustre: llite: limit max xattr size by kernel value
lustre: llite: report latency for filesystem ops
lustre: osc: allow increasing osc.*.short_io_bytes
lustre: ptlrpc: update wiretest for new values
lustre: uapi: LU-12521 llapi: add separate fsname and instance API
lustre: ptlrpc: show target name in req_history
lustre: llite: proper names/types for offset/pages
lustre: uapi: remove unused LUSTRE_DIRECTIO_FL
lnet: use conservative health timeouts
Andrew Perepechko (5):
lustre: build: armv7 client build fixes
lustre: osc: speed up page cache cleanup during blocking ASTs
lustre: ptlrpc: improve memory allocation for service RPCs
lustre: llite: optimizations for not granted lock processing
lnet: libcfs: crashes with certain cpu part numbers
Andriy Skulysh (16):
lustre: ptlrpc: ptlrpc_register_bulk() LBUG on ENOMEM
lustre: ptlrpc: Serialize procfs access to scp_hist_reqs using mutex
lustre: ptlrpc: ASSERTION(!list_empty(imp->imp_replay_cursor))
lustre: ptlrpc: connect vs import invalidate race
lnet: o2iblnd: ibc_rxs is created and freed with different size
lustre: ldlm: Lost lease lock on migrate error
lnet: o2iblnd: kib_conn leak
lustre: ptlrpc: Bulk assertion fails on -ENOMEM
lustre: ptlrpc: ASSERTION (req_transno < next_transno) failed
lustre: ptlrpc: ocd_connect_flags are wrong during reconnect
lustre: ptlrpc: Add increasing XIDs CONNECT2 flag
lustre: ptlrpc: don't reset lru_resize on idle reconnect
lustre: ptlrpc: resend may corrupt the data
lustre: ldlm: FLOCK request can be processed twice
lustre: ldlm: signal vs CP callback race
lustre: llite: eviction during ll_open_cleanup()
Ann Koehler (8):
lustre: llite: yield cpu after call to ll_agl_trigger
lustre: ptlrpc: Do not map unrecognized ELDLM errnos to EIO
lustre: llite: Lock inode on tiny write if setuid/setgid set
lustre: statahead: sa_handle_callback get lli_sa_lock earlier
lustre: ptlrpc: Add jobid to rpctrace debug messages
lnet: libcfs: Reduce memory frag due to HA debug msg
lustre: llite: release active extent on sync write commit
lustre: ptlrpc: ptlrpc_register_bulk LBUG on ENOMEM
Arshad Hussain (17):
lustre: osc: truncate does not update blocks count on client
lustre: lmv: Fix style issues for lmv_fld.c
lustre: llite: Fix style issues for llite_nfs.c
lustre: llite: Fix style issues for lcommon_misc.c
lustre: llite: Fix style issues for symlink.c
lustre: ptlrpc: Change static defines to use macro for sec_gc.c
lustre: ldlm: Fix style issues for ldlm_lockd.c
lustre: ldlm: Fix style issues for ldlm_request.c
lustre: ptlrpc: Fix style issues for sec_bulk.c
lustre: ldlm: Fix style issues for ptlrpcd.c
lustre: ptlrpc: Fix style issues for sec_null.c
lustre: ptlrpc: Fix style issues for service.c
lustre: ldlm: Fix style issues for ldlm_resource.c
lustre: ptlrpc: Fix style issues for sec_gc.c
lustre: ptlrpc: Fix style issues for llog_client.c
lnet: Change static defines to use macro for module.c
lustre: ldlm: Fix style issues for ldlm_lib.c
Artem Blagodarenko (1):
lnet: add fault injection for bulk transfers
Aurelien Degremont (1):
lnet: support non-default network namespace
Ben Evans (1):
lustre: headers: define pct(a,b) once
Bobi Jam (11):
lustre: osc: depart grant shrinking from pinger
lustre: osc: enable/disable OSC grant shrink
lustre: flr: add 'nosync' flag for FLR mirrors
lustre: mdc: grow lvb buffer to hold layout
lustre: flr: add mirror write command
lustre: llite: protect reading inode->i_data.nrpages
lustre: osc: limit chunk number of write submit
lustre: osc: prevent use after free
lustre: llite: error handling of ll_och_fill()
lustre: flr: avoid reading unhealthy mirror
lustre: llite: file write pos mimatch
Bruno Faccini (6):
lustre: obdclass: fix llog_cat_cleanup() usage on Client
lustre: ptlrpc: fix test_req_buffer_pressure behavior
lustre: ldlm: cleanup LVB handling
lustre: security: return security context for metadata ops
lustre: lov: new foreign LOV format
lustre: lmv: new foreign LMV format
Chris Horn (28):
lnet: Cleanup lnet_get_rtr_pool_cfg
lnet: Fix NI status in debugfs for loopback ni
lustre: ptlrpc: Add more flags to DEBUG_REQ_FLAGS macro
lnet: Protect lp_dc_pendq manipulation with lp_lock
lnet: Ensure md is detached when msg is not committed
lnet: Do not allow gateways on remote nets
lnet: Convert noisy timeout error to cdebug
lnet: Misleading error from lnet_is_health_check
lnet: Sync the start of discovery and monitor threads
lnet: Deprecate live and dead router check params
lnet: Detach rspt when md_threshold is infinite
lnet: Return EHOSTUNREACH for unreachable gateway
lnet: Defer rspt cleanup when MD queued for unlink
lnet: Don't queue msg when discovery has completed
lnet: Use alternate ping processing for non-mr peers
lnet: o2ib: Record rc in debug log on startup failure
lnet: o2ib: Reintroduce kiblnd_dev_search
lnet: Optimize check for routing feature flag
lnet: Wait for single discovery attempt of routers
lnet: Prefer route specified by rtr_nid
lnet: Add peer level aliveness information
lnet: Refactor lnet_find_best_lpni_on_net
lnet: Avoid comparing route to itself
lnet: Avoid extra lnet_remotenet lookup
lnet: Remove unused vars in lnet_find_route_locked
lnet: Refactor lnet_compare_routes
lnet: Fix source specified route selection
lnet: Do not assume peers are MR capable
Christopher J. Morrone (1):
lustre: ldlm: Make kvzalloc | kvfree use consistent
Di Wang (1):
lustre: llite: handle ORPHAN/DEAD directories
Emoly Liu (4):
lnet: fix nid range format '*@' support
lustre: checksum: enable/disable checksum correctly
lustre: ptlrpc: check lm_bufcount and lm_buflen
lustre: ptlrpc: check buffer length in lustre_msg_string()
Fan Yong (3):
lustre: llite: return compatible fsid for statfs
lustre: llite: decrease sa_running if fail to start statahead
lustre: lfsck: layout LFSCK for mirrored file
Gu Zheng (3):
lustre: osc: cancel osc_lock list traversal once found the lock is
being used
lustre: ldlm: always cancel aged locks regardless enabling or
disabling lru resize
lustre: uapi: fix building fail against Power9 little endian
Hongchao Zhang (8):
lustre: mdc: resend quotactl if needed
lustre: quota: add default quota setting support
lustre: ptlrpc: race in AT early reply
lustre: ptlrpc: always unregister bulk
lustre: quota: protect quota flags at OSC
lustre: quota: make overquota flag for old req
lustre: fld: let's caller to retry FLD_QUERY
lustre: mdc: hold obd while processing changelog
Jacek Tomaka (1):
lustre: llite: Mark lustre_inode_cache as reclaimable
Jadhav Vikram (1):
lustre: lov: protected ost pool count updation
James Nunez (1):
lustre: llite: limit smallest max_cached_mb value
James Simmons (33):
lustre: always enable special debugging, fhandles, and quota support.
lustre: osc_cache: remove __might_sleep()
lustre: uapi: remove enum hsm_progress_states
lustre: uapi: sync enum obd_statfs_state
lustre: obd: create ping sysfs file
lustre: ldlm: change LDLM_POOL_ADD_VAR macro to inline function
lustre: osc: fix idle_timeout handling
lustre: ptlrpc: replace simple_strtol with kstrtol
lustre: obd: use correct ip_compute_csum() version
lustre: llite: create checksums to replace checksum_pages
lustre: obd: use correct names for conn_uuid
lustre: mgc: restore mgc binding for sptlrpc
lustre: update version to 2.11.99
lustre: obdclass: report all obd states for OBD_IOC_GETDEVICE
lustre: sysfs: make ping sysfs file read and writable
lustre: sptlrpc: split sptlrpc_process_config()
lustre: clio: fix incorrect invariant in cl_io_iter_fini()
lustre: obd: use ldo_process_config for mdc and osc layer
lustre: obd: make health_check sysfs compliant
lnet: properly cleanup lnet debugfs files
lustre: obd: update udev event handling
lustre: obd: replace class_uuid with linux kernel version.
lustre: obd: round values to nearest MiB for *_mb syfs files
lustre: obdclass: add comment for rcu handling in lu_env_remove
lustre: ptlrpc: change IMPORT_SET_* macros into real functions
lustre: obd: harden debugfs handling
lustre: update version to 2.13.50
lnet: timers: correctly offset mod_timer.
lustre: obd: perform proper division
lustre: llite: don't cache MDS_OPEN_LOCK for volatile files
lnet: socklnd: rename struct ksock_peer to struct ksock_peer_ni
lustre: sysfs: use string helper like functions for sysfs
lustre: uapi: properly pack data structures
Jian Yu (4):
lustre: mdt: revoke lease lock for truncate
lustre: llite: swab LOV EA user data
lustre: llite: swab LOV EA data in ll_getxattr_lov()
lustre: llite: fetch default layout for a directory
Jinshan Xiong (4):
lustre: llite: rename FSFILT_IOC_* to system flags
lustre: llite: optimize read on open pages
lustre: dne: performance improvement for file creation
lustre: llite: do not cache write open lock for exec file
John L. Hammond (12):
lustre: llite: reorganize variable and data structures
lustre: hsm: ignore compound_id
lustre: llog: remove obsolete llog handlers
lustre: obd: keep dirty_max_pages a round number of MB
lustre: llite: handle zero length xattr values correctly
lustre: mdc: remove obsolete intent opcodes
lustre: ldlm: correct logic in ldlm_prepare_lru_list()
lustre: llite: zero lum for stripeless files
lustre: mdc: move empty xattr handling to mdc layer
lustre: obd: remove portals handle from OBD import
lustre: llite: handle -ENODATA in ll_layout_fetch()
lustre: ldlm: remove trace from ldlm_pool_count()
Kit Westneat (1):
lnet: remove .nf_min_max handling
Lai Siyao (27):
lustre: ptlrpc: add dir migration connect flag
lustre: lmv: dir page is released while in use
lustre: migrate: pack lmv ea in migrate rpc
lustre: migrate: migrate striped directory
lustre: lmv: support accessing migrating directory
lustre: llite: add lock for dir layout data
lustre: lmv: allocate fid on parent MDT in migrate
lustre: obdclass: lu_dirent record length missing '0'
lustre: uapi: reserve connect flag for plain layout
lustre: dne: allow access to striped dir with broken layout
lustre: dne: add new dir hash type "space"
lustre: ptlrpc: intent_getattr fetches default LMV
lustre: mdc: add async statfs
lustre: lmv: mkdir with balanced space usage
lustre: lmv: reuse object alloc QoS code from LOD
lustre: obdclass: generate random u64 max correctly
lustre: uapi: change "space" hash type to hash flag
lustre: obdclass: 0-nlink race in lu_object_find_at()
lustre: mdc: dir page ldp_hash_end mistakenly adjusted
lustre: lmv: disable remote file statahead
lustre: lmv: use lu_tgt_descs to manage tgts
lustre: lmv: share object alloc QoS code with LMV
lustre: obdclass: qos penalties miscalculated
lustre: obdclass: lu_tgt_descs cleanup
lustre: lmv: alloc dir stripes by QoS
lustre: uapi: introduce OBD_CONNECT2_CRUSH
lustre: llite: fix deadlock in ll_update_lsm_md()
Li Dongyang (9):
lustre: ldlm: check double grant race after resource change
lustre: clio: use pagevec_release for many pages
lustre: osc: reduce atomic ops in osc_enter_cache_try
lustre: osc: check if opg is in lru list without locking
lustre: osc: don't check capability for every page
lustre: osc: reduce lock contention in osc_unreserve_grant
lustre: obdclass: protect imp_sec using rwlock_t
lustre: llite: create obd_device with usercopy whitelist
lustre: obdclass: remove assertion for imp_refcount
Li Xi (6):
lustre: osc: add T10PI support for RPC checksum
lustre: osc: wrong page offset for T10PI checksum
lustre: llite: add file heat support
lustre: llite: console message for disabled flock call
lustre: llite: cleanup stats of LPROC_LL_*
lustre: osc: add preferred checksum type support
Liang Zhen (2):
lustre: ldlm: don't disable softirq for exp_rpc_lock
lustre: obdclass: new wrapper to convert NID to string
Mike Marciniszyn (1):
lnet: libcfs: remove unnecessary set_fs(KERNEL_DS)
Mikhail Pershin (28):
lustre: mdc: deny layout swap for DoM file
lustre: ldlm: expose dirty age limit for flush-on-glimpse
lustre: ldlm: IBITS lock convert instead of cancel
lustre: ptlrpc: add LOCK_CONVERT connection flag
lustre: ldlm: handle lock converts in cancel handler
lustre: ldlm: don't add canceling lock back to LRU
lustre: mdt: read on open for DoM files
lustre: llite: check truncate race for DOM pages
lustre: ldlm: don't cancel DoM locks before replay
lustre: ptlrpc: don't change buffer when signature is ready
lustre: ldlm: update l_blocking_lock under lock
lustre: ldlm: don't apply ELC to converting and DOM locks
lustre: ldlm: don't skip bl_ast for local lock
lustre: mdt: fix read-on-open for big PAGE_SIZE
lustre: ldlm: don't convert wrong resource
lustre: mdc: prevent glimpse lock count grow
lustre: mdc: return DOM size on open resend
lustre: osc: pass client page size during reconnect too
lustre: mdt: fix mdt_dom_discard_data() timeouts
lustre: dom: per-resource ELC for WRITE lock enqueue
lustre: dom: mdc_lock_flush() improvement
lustre: obdclass: remove unprotected access to lu_object
lustre: llite: check correct size in ll_dom_finish_open()
lustre: ptlrpc: fix reply buffers shrinking and growing
lustre: dom: manual OST-to-DOM migration via mirroring
lustre: ptlrpc: do lu_env_refill for any new request
lustre: dom: check read-on-open buffer presents in reply
lustre: llog: keep llog handle alive until last reference
Mr NeilBrown (47):
lustre: obdclass: allow per-session jobids.
lustre: fld: remove fci_no_shrink field.
lustre: lustre: remove ldt_obd_type field of lu_device_type
lustre: lustre: remove imp_no_timeout field
lustre: llog: remove olg_cat_processing field.
lustre: ptlrpc: remove struct ptlrpc_bulk_page
lustre: ptlrpc: remove bd_import_generation field.
lustre: ptlrpc: remove srv_threads from struct ptlrpc_service
lustre: ptlrpc: remove scp_nthrs_stopping field.
lustre: ldlm: remove unused ldlm_server_conn
lustre: llite: remove lli_readdir_mutex
lustre: llite: remove ll_umounting field
lustre: llite: align field names in ll_sb_info
lustre: llite: remove lti_iter field
lustre: llite: remove ft_mtime field
lustre: llite: remove sub_reenter field.
lustre: osc: remove oti_descr oti_handle oti_plist
lustre: osc: remove oe_next_page
lnet: o2iblnd: remove some unused fields.
lnet: socklnd: remove ksnp_sharecount
lnet: change ln_mt_waitq to a completion.
lustre: import: Fix missing spin_unlock()
lustre: use simple sleep in some cases
lustre: modules: Use LIST_HEAD for declaring list_heads
lnet: remove pt_number from lnet_peer_table.
lustre: obdclass: Allow read-ahead for write requests
lnet: discard lnd_refcount
lnet: change ksocknal_create_peer() to return pointer
lnet: discard ksnn_lock
lnet: discard LNetMEInsert
lustre: all: prefer sizeof(*var) for alloc
lnet: always check return of try_module_get()
lnet: prepare to make lnet_lnd const.
lnet: discard struct ksock_peer
lnet: socklnd: initialize the_ksocklnd at compile-time.
lnet: remove locking protection ln_testprotocompat
lustre: handle: remove locking from class_handle2object()
lustre: obdclass: convert waiting in cl_sync_io_wait().
lnet: modules: use list_move were appropriate.
lnet: fix small race in unloading klnd modules.
lnet: me: discard struct lnet_handle_me
lnet: socklnd: convert peers hash table to hashtable.h
lustre: ptlrpc: simplify wait_event handling in unregister functions
lustre: ptlrpc: use l_wait_event_abortable in ptlrpcd_add_reg()
lnet: use LIST_HEAD() for local lists.
lustre: lustre: use LIST_HEAD() for local lists.
lnet: remove lnd_query interface.
Nathaniel Clark (1):
lustre: lov: Correct bounds checking
NeilBrown (18):
lustre: llite: Don't clear d_fsdata in ll_release()
lustre: llite: move agl_thread cleanup out of thread.
lustre/lnet: remove unnecessary use of msecs_to_jiffies()
lnet: net_fault: don't pass struct member to do_div()
lustre: obd: discard unused enum
lustre: lov: use wait_event() in lov_subobject_kill()
lustre: llite: use wait_event in cl_object_put_last()
lustre: handle: move refcount into the lustre_handle.
lustre: ldlm: separate buckets from ldlm hash table
lustre: handle: discard OBD_FREE_RCU
lnet: use list_move where appropriate.
lustre: ldlm: add a counter to the per-namespace data
lustre: rename ops to owner
lustre: ldlm: simplify ldlm_ns_hash_defs[]
lustre: u_object: factor out extra per-bucket data
lustre: llite: replace lli_trunc_sem
lustre: handle: use hlist for hash lists.
lustre: handle: discard h_lock.
Olaf Faaland (2):
lnet: create existing net returns EEXIST
lustre: llite: Update mdc and lite stats on open|creat
Olaf Weber (1):
lnet: use after free in lnet_discover_peer_locked()
Oleg Drokin (6):
lustre: ptlrpc: Add WBC connect flag
lustre: lov: Move lov_tgts_kobj init to lov_setup
lustre: osc: increase default max_dirty_mb to 2G
lustre: llite: Revalidate dentries in ll_intent_file_open
lustre: llite: hash just created files if lock allows
lustre: ptlrpc: Properly swab ll_fiemap_info_key
Patrick Farrell (30):
lustre: osc: Do not request more than 2GiB grant
lustre: ldlm: Reduce debug to console during eviction
lustre: ptlrpc: Make CPU binding switchable
lustre: osc: Do not walk full extent list
lustre: ldlm: Adjust search_* functions
lustre: mdc: Improve xattr buffer allocations
lustre: llite: Initialize cl_dirty_max_pages
lustre: llite: ll_fault fixes
lustre: osd: Set max ea size to XATTR_SIZE_MAX
lustre: lov: Remove unnecessary assert
lustre: obd: Add overstriping CONNECT flag
lustre: lov: Add overstriping support
lustre: uapi: Add nonrotational flag to statfs
lustre: llite: collect debug info for ll_fsync
lustre: lu_object: Add missed qos_rr_init
lustre: osc: Do not assert for first extent
lustre: ptlrpc: Don't get jobid in body_v2
lustre: lov: Correct write_intent end for trunc
lustre: osc: Fix dom handling in weight_ast
lustre: llite: Fix extents_stats
lustre: ptlrpc: Stop sending ptlrpc_body_v2
lustre: uapi: Remove unused CONNECT flag
lustre: llite: Fix page count for unaligned reads
lustre: llite: Improve readahead RPC issuance
lustre: lov: Move page index to top level
lustre: ptlrpc: Hold imp lock for idle reconnect
lustre: osc: glimpse - search for active lock
lnet: o2iblnd: Make credits hiw connection aware
lustre: vvp: dirty pages with pagevec
lustre: llite: Accept EBUSY for page unaligned read
Qian Yingjin (16):
lustre: mdt: Lazy size on MDT
lustre: uapi: add new changerec_type
lustre: lsom: Add an OBD_CONNECT2_LSOM connect flag
lustre: pcc: Reserve a new connection flag for PCC
lustre: rpc: support maximum 64MB I/O RPC
lustre: llite: Add persistent cache on client
lustre: pcc: Non-blocking PCC caching
lustre: pcc: security and permission for non-root user access
lustre: llite: Rule based auto PCC caching when create files
lustre: pcc: auto attach during open for valid cache
lustre: pcc: change detach behavior and add keep option
lustre: som: integrate LSOM with lfs find
lustre: pcc: Auto attach for PCC during IO
lustre: pcc: Incorrect size after re-attach
lustre: pcc: auto attach not work after client cache clear
lustre: pcc: Init saved dataset flags properly
Quentin Bouget (1):
lustre: uapi: turn struct lustre_nfs_fid to userland fhandle
Rahul Deshmukh (1):
lustre: obdecho: turn on async flag only for mode 3
Rob Latham (1):
lustre: uapi: Make lustre_user.h c++-legal
Ryan Haasken (1):
lustre: obdclass: Add lbug_on_eviction option
Sebastien Buisson (7):
lustre: obd: check '-o network' and peer discovery conflict
lustre: cfg: reserve flags for SELinux status checking
lustre: sec: create new function sptlrpc_get_sepol()
lnet: check for asymmetrical route messages
lustre: ptlrpc: manage SELinux policy info at connect time
lustre: ptlrpc: manage SELinux policy info for metadata ops
lustre: sec: reserve flags for client side encryption
Sergey Cheremencev (1):
lustre: ptlrpc: IR doesn't reconnect after EAGAIN
Shaun Tancheff (7):
lustre: lov: return error if cl_env_get fails
lustre: llite: MS_* flags and SB_* flags split
lustre: clio: support custom csi_end_io handler
lnet: Fix style issues for selftest/rpc.c
lnet: Fix style issues for module.c conctl.c
lnet: libcfs: provide an scnprintf and start using it
lnet: libcfs: Cleanup use of bare printk
Sonia Sharma (6):
lnet: Fix selftest backward compatibility post health
lnet: socklnd: dynamically set LND parameters
lnet: peer deletion code may hide error
lnet: increase lnet transaction timeout
lnet: Avoid lnet debugfs read/write if ctl_table does not exist
lnet: check if current->nsproxy is NULL before using
Swapnil Pimpale (1):
lustre: lustre: Reserve OST_FALLOCATE(fallocate) opcode
Tatsushi Takamura (1):
lnet: handling device failure by IB event handler
Teddy Chan (1):
lustre: ptlrpc: Add QoS for uid and gid in NRS-TBF
Teddy Zheng (2):
lustre: hsm: add OBD_CONNECT2_ARCHIVE_ID_ARRAY to pass archive_id
lists in array
lustre: hsm: increase upper limit of maximum HSM backends registered
with MDT
Vitaly Fertman (6):
lustre: ptlrpc: Add more flags to DEBUG_REQ_FLAGS macro
lustre: ldlm: layout lock fixes
lustre: osc: layout and chunkbits alignment mismatch
lustre: osc: wrong cache of LVB attrs
lustre: osc: wrong cache of LVB attrs, part2
lustre: ldlm: fix lock convert races
Vladimir Saveliev (6):
lustre: obdclass: make mod rpc slot wait queue FIFO
lustre: lov: fix lov_iocontrol for inactive OST case
lnet: libcfs: do not calculate debug_mb if it is set
lustre: llite: improve ll_dom_lock_cancel
lustre: lov: check all entries in lov_flush_composite
lustre: lmv: disable statahead for remote objects
Wang Shilong (24):
lustre: llite: fix setstripe for specific osts upon dir
lnet: libcfs: fix wrong check in libcfs_debug_vmsg2()
lustre: quota: fix setattr project check
lustre: llite: make sure name pack atomic
lustre: ptlrpc: handle proper import states for recovery
lustre: llite: switch to use ll_fsname directly
lustre: llite: fill copied dentry name's ending char properly
lustre: llite, readahead: fix to call ll_ras_enter() properly
lnet: libcfs: fix panic for too large cpu partitions
lustre: lov: fix wrong calculated length for fiemap
lustre: push rcu_barrier() before destroying slab
lustre: llite,readahead: don't always use max RPC size
lustre: llite: improve single-thread read performance
lustre: llite: fix deadloop with tiny write
lustre: llite: make sure readahead cover current read
lustre: llite: don't check vmpage refcount in ll_releasepage()
lustre: osc: reserve lru pages for read in batch
lustre: llite: don't miss every first stride page
lustre: llite: extend readahead locks for striped file
lustre: readahead: convert stride page index to byte
lnet: eliminate uninitialized warning
lustre: llite: support page unaligned stride readahead
lustre: ptlrpc: always reset generation for idle reconnect
lustre: lmv: fix to return correct MDT count
Yang Sheng (6):
lustre: ldlm: speed up preparation for list of lock cancel
lustre: ldlm: fix for l_lru usage
lustre: class: use INIT_LIST_HEAD_RCU instead INIT_LIST_HEAD
lustre: lov: cl_cache could miss initialize
lustre: lov: remove KEY_CACHE_SET to simplify the code
lustre: import: fix race between imp_state & imp_invalid
fs/lustre/Kconfig | 5 +
fs/lustre/fid/fid_request.c | 7 +
fs/lustre/fld/fld_cache.c | 15 +-
fs/lustre/fld/fld_internal.h | 1 -
fs/lustre/fld/fld_request.c | 23 +-
fs/lustre/include/cl_object.h | 73 +-
fs/lustre/include/lprocfs_status.h | 33 +-
fs/lustre/include/lu_object.h | 238 +-
fs/lustre/include/lustre_disk.h | 1 +
fs/lustre/include/lustre_dlm.h | 155 +-
fs/lustre/include/lustre_dlm_flags.h | 33 +-
fs/lustre/include/lustre_export.h | 29 +-
fs/lustre/include/lustre_ha.h | 2 +-
fs/lustre/include/lustre_handles.h | 21 +-
fs/lustre/include/lustre_import.h | 45 +-
fs/lustre/include/lustre_lmv.h | 82 +-
fs/lustre/include/lustre_log.h | 4 +-
fs/lustre/include/lustre_mdc.h | 120 -
fs/lustre/include/lustre_net.h | 159 +-
fs/lustre/include/lustre_osc.h | 27 +-
fs/lustre/include/lustre_req_layout.h | 15 +-
fs/lustre/include/lustre_sec.h | 12 +
fs/lustre/include/lustre_swab.h | 2 +
fs/lustre/include/obd.h | 151 +-
fs/lustre/include/obd_cksum.h | 130 +-
fs/lustre/include/obd_class.h | 150 +-
fs/lustre/include/obd_support.h | 59 +-
fs/lustre/ldlm/ldlm_extent.c | 2 +-
fs/lustre/ldlm/ldlm_inodebits.c | 111 +-
fs/lustre/ldlm/ldlm_internal.h | 42 +-
fs/lustre/ldlm/ldlm_lib.c | 59 +-
fs/lustre/ldlm/ldlm_lock.c | 359 +--
fs/lustre/ldlm/ldlm_lockd.c | 224 +-
fs/lustre/ldlm/ldlm_pool.c | 28 +-
fs/lustre/ldlm/ldlm_request.c | 444 +++-
fs/lustre/ldlm/ldlm_resource.c | 196 +-
fs/lustre/llite/Makefile | 2 +-
fs/lustre/llite/dcache.c | 1 -
fs/lustre/llite/dir.c | 731 ++++--
fs/lustre/llite/file.c | 987 ++++++--
fs/lustre/llite/glimpse.c | 1 +
fs/lustre/llite/lcommon_cl.c | 35 +-
fs/lustre/llite/lcommon_misc.c | 47 +-
fs/lustre/llite/llite_internal.h | 444 +++-
fs/lustre/llite/llite_lib.c | 719 ++++--
fs/lustre/llite/llite_mmap.c | 63 +-
fs/lustre/llite/llite_nfs.c | 59 +-
fs/lustre/llite/lproc_llite.c | 465 +++-
fs/lustre/llite/namei.c | 717 ++++--
fs/lustre/llite/pcc.c | 2614 ++++++++++++++++++++
fs/lustre/llite/pcc.h | 264 ++
fs/lustre/llite/rw.c | 1091 ++++++---
fs/lustre/llite/rw26.c | 4 -
fs/lustre/llite/statahead.c | 182 +-
fs/lustre/llite/super25.c | 23 +-
fs/lustre/llite/symlink.c | 21 +-
fs/lustre/llite/vvp_dev.c | 1 +
fs/lustre/llite/vvp_internal.h | 22 +-
fs/lustre/llite/vvp_io.c | 176 +-
fs/lustre/llite/vvp_object.c | 11 +-
fs/lustre/llite/vvp_page.c | 19 +-
fs/lustre/llite/xattr.c | 109 +-
fs/lustre/llite/xattr_security.c | 19 +
fs/lustre/lmv/lmv_fld.c | 17 +-
fs/lustre/lmv/lmv_intent.c | 201 +-
fs/lustre/lmv/lmv_internal.h | 162 +-
fs/lustre/lmv/lmv_obd.c | 2078 +++++++++-------
fs/lustre/lmv/lproc_lmv.c | 143 +-
fs/lustre/lov/Makefile | 2 +-
fs/lustre/lov/lov_cl_internal.h | 28 +-
fs/lustre/lov/lov_ea.c | 117 +-
fs/lustre/lov/lov_internal.h | 49 +-
fs/lustre/lov/lov_io.c | 89 +-
fs/lustre/lov/lov_obd.c | 162 +-
fs/lustre/lov/lov_object.c | 159 +-
fs/lustre/lov/lov_offset.c | 2 +
fs/lustre/lov/lov_pack.c | 73 +-
fs/lustre/lov/lov_page.c | 17 +-
fs/lustre/lov/lov_pool.c | 19 +-
fs/lustre/lov/lov_request.c | 29 +-
fs/lustre/lov/lovsub_page.c | 68 -
fs/lustre/lov/lproc_lov.c | 4 +-
fs/lustre/mdc/lproc_mdc.c | 87 +-
fs/lustre/mdc/mdc_changelog.c | 154 +-
fs/lustre/mdc/mdc_dev.c | 171 +-
fs/lustre/mdc/mdc_internal.h | 14 +-
fs/lustre/mdc/mdc_lib.c | 86 +-
fs/lustre/mdc/mdc_locks.c | 277 ++-
fs/lustre/mdc/mdc_reint.c | 106 +-
fs/lustre/mdc/mdc_request.c | 476 +++-
fs/lustre/mgc/lproc_mgc.c | 12 +-
fs/lustre/mgc/mgc_request.c | 86 +-
fs/lustre/obdclass/Makefile | 2 +-
fs/lustre/obdclass/cl_io.c | 49 +-
fs/lustre/obdclass/cl_object.c | 23 +-
fs/lustre/obdclass/cl_page.c | 36 +-
fs/lustre/obdclass/class_obd.c | 151 +-
fs/lustre/obdclass/genops.c | 139 +-
fs/lustre/obdclass/integrity.c | 273 +++
fs/lustre/obdclass/jobid.c | 282 ++-
fs/lustre/obdclass/llog.c | 126 +-
fs/lustre/obdclass/llog_cat.c | 59 +-
fs/lustre/obdclass/llog_internal.h | 4 +-
fs/lustre/obdclass/lprocfs_status.c | 277 ++-
fs/lustre/obdclass/lu_object.c | 518 +++-
fs/lustre/obdclass/lu_tgt_descs.c | 682 ++++++
fs/lustre/obdclass/lustre_handles.c | 61 +-
fs/lustre/obdclass/obd_cksum.c | 151 ++
fs/lustre/obdclass/obd_config.c | 39 +-
fs/lustre/obdclass/obd_mount.c | 23 +-
fs/lustre/obdclass/obd_sysfs.c | 101 +-
fs/lustre/obdclass/obdo.c | 7 +-
fs/lustre/obdecho/echo_client.c | 77 +-
fs/lustre/osc/lproc_osc.c | 232 +-
fs/lustre/osc/osc_cache.c | 172 +-
fs/lustre/osc/osc_dev.c | 19 +-
fs/lustre/osc/osc_internal.h | 48 +-
fs/lustre/osc/osc_io.c | 115 +-
fs/lustre/osc/osc_lock.c | 156 +-
fs/lustre/osc/osc_object.c | 28 +-
fs/lustre/osc/osc_page.c | 20 +-
fs/lustre/osc/osc_quota.c | 18 +-
fs/lustre/osc/osc_request.c | 619 +++--
fs/lustre/ptlrpc/client.c | 282 ++-
fs/lustre/ptlrpc/errno.c | 27 +
fs/lustre/ptlrpc/events.c | 12 +-
fs/lustre/ptlrpc/import.c | 507 ++--
fs/lustre/ptlrpc/layout.c | 172 +-
fs/lustre/ptlrpc/llog_client.c | 15 +-
fs/lustre/ptlrpc/lproc_ptlrpc.c | 66 +-
fs/lustre/ptlrpc/niobuf.c | 102 +-
fs/lustre/ptlrpc/pack_generic.c | 236 +-
fs/lustre/ptlrpc/pinger.c | 50 +-
fs/lustre/ptlrpc/ptlrpc_internal.h | 3 +-
fs/lustre/ptlrpc/ptlrpcd.c | 21 +-
fs/lustre/ptlrpc/recover.c | 23 +-
fs/lustre/ptlrpc/sec.c | 146 +-
fs/lustre/ptlrpc/sec_bulk.c | 71 +-
fs/lustre/ptlrpc/sec_config.c | 89 +-
fs/lustre/ptlrpc/sec_gc.c | 16 +-
fs/lustre/ptlrpc/sec_lproc.c | 74 +
fs/lustre/ptlrpc/sec_null.c | 16 +-
fs/lustre/ptlrpc/sec_plain.c | 7 +-
fs/lustre/ptlrpc/service.c | 427 ++--
fs/lustre/ptlrpc/wiretest.c | 342 ++-
include/linux/libcfs/libcfs.h | 1 +
include/linux/libcfs/libcfs_debug.h | 69 +-
include/linux/libcfs/libcfs_fail.h | 46 +-
include/linux/lnet/api.h | 34 +-
include/linux/lnet/lib-lnet.h | 225 +-
include/linux/lnet/lib-types.h | 355 ++-
include/uapi/linux/lnet/libcfs_debug.h | 4 +-
include/uapi/linux/lnet/libcfs_ioctl.h | 13 +-
include/uapi/linux/lnet/lnet-dlc.h | 42 +
include/uapi/linux/lnet/lnet-types.h | 49 +-
include/uapi/linux/lnet/lnetctl.h | 23 +
include/uapi/linux/lnet/nidstr.h | 2 +
include/uapi/linux/lustre/lustre_cfg.h | 1 +
include/uapi/linux/lustre/lustre_fid.h | 7 +
include/uapi/linux/lustre/lustre_idl.h | 392 +--
include/uapi/linux/lustre/lustre_ioctl.h | 5 +-
include/uapi/linux/lustre/lustre_kernelcomm.h | 15 +-
include/uapi/linux/lustre/lustre_user.h | 575 +++--
include/uapi/linux/lustre/lustre_ver.h | 6 +-
mm/page-writeback.c | 1 +
net/lnet/klnds/o2iblnd/o2iblnd.c | 357 ++-
net/lnet/klnds/o2iblnd/o2iblnd.h | 63 +-
net/lnet/klnds/o2iblnd/o2iblnd_cb.c | 169 +-
net/lnet/klnds/o2iblnd/o2iblnd_modparams.c | 30 +-
net/lnet/klnds/socklnd/socklnd.c | 737 +++---
net/lnet/klnds/socklnd/socklnd.h | 95 +-
net/lnet/klnds/socklnd/socklnd_cb.c | 139 +-
net/lnet/klnds/socklnd/socklnd_proto.c | 24 +-
net/lnet/libcfs/debug.c | 5 +-
net/lnet/libcfs/fail.c | 15 +-
net/lnet/libcfs/libcfs_cpu.c | 11 +-
net/lnet/libcfs/libcfs_lock.c | 2 +-
net/lnet/libcfs/linux-crypto.c | 5 +-
net/lnet/libcfs/module.c | 33 +-
net/lnet/libcfs/tracefile.c | 50 +-
net/lnet/lnet/acceptor.c | 27 +-
net/lnet/lnet/api-ni.c | 828 +++++--
net/lnet/lnet/config.c | 57 +-
net/lnet/lnet/lib-eq.c | 4 +-
net/lnet/lnet/lib-md.c | 23 +-
net/lnet/lnet/lib-me.c | 135 +-
net/lnet/lnet/lib-move.c | 3266 +++++++++++++++++++------
net/lnet/lnet/lib-msg.c | 711 +++++-
net/lnet/lnet/lib-ptl.c | 2 +-
net/lnet/lnet/lib-socket.c | 17 +-
net/lnet/lnet/lo.c | 1 -
net/lnet/lnet/module.c | 8 +-
net/lnet/lnet/net_fault.c | 126 +-
net/lnet/lnet/nidstrings.c | 272 +-
net/lnet/lnet/peer.c | 725 ++++--
net/lnet/lnet/router.c | 1613 ++++++------
net/lnet/lnet/router_proc.c | 212 +-
net/lnet/selftest/conctl.c | 4 +-
net/lnet/selftest/console.c | 10 +-
net/lnet/selftest/framework.c | 28 +-
net/lnet/selftest/module.c | 2 +-
net/lnet/selftest/rpc.c | 43 +-
net/lnet/selftest/rpc.h | 10 +-
203 files changed, 25222 insertions(+), 10485 deletions(-)
create mode 100644 fs/lustre/llite/pcc.c
create mode 100644 fs/lustre/llite/pcc.h
delete mode 100644 fs/lustre/lov/lovsub_page.c
create mode 100644 fs/lustre/obdclass/integrity.c
create mode 100644 fs/lustre/obdclass/lu_tgt_descs.c
create mode 100644 fs/lustre/obdclass/obd_cksum.c
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:50 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:50 -0500
Subject: [lustre-devel] [PATCH 002/622] lustre: osc_cache: remove
__might_sleep()
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-3-git-send-email-jsimmons@infradead.org>
The patch 'simplify osc_wake_cache_waiters()' created a new
wrapper wait_event_idle_exclusive_timeout_cmd() which includes
a __might_sleep() test. This was causing the following back
trace:
kernel: BUG: sleeping function called from invalid context at fs/lustre/osc/osc_cache.c:1635
kernel: in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 19374, name: cp
kernel: INFO: lockdep is turned off.
kernel: Preemption disabled at:
kernel: [<0000000000000000>] 0x0
kernel: CPU: 11 PID: 19374 Comm: cp Tainted: G W 5.4.0-rc5+ #1
kernel: Call Trace:
kernel: dump_stack+0x5e/0x8b
kernel: ___might_sleep+0x205/0x260
kernel: osc_queue_async_io+0x1104/0x1de0 [osc]
kernel: ? _raw_spin_unlock+0x2e/0x50
kernel: ? libcfs_debug_msg+0x6ab/0xc80 [libcfs]
kernel: ? vvp_io_setattr_start+0x200/0x200 [lustre]
kernel: osc_page_cache_add+0x2c/0xa0 [osc]
kernel: osc_io_commit_async+0x1a8/0x420 [osc]
kernel: cl_io_commit_async+0x58/0x80 [obdclass]
kernel: ? vvp_io_setattr_start+0x200/0x200 [lustre:1
This can be called from an atomic context and examing the code
suggest we don't need __might_sleep() so lets remove it.
Fixes: def8e96d4f3d ("lustre: osc_cache: simplify osc_wake_cache_waiters()")
Signed-off-by: James Simmons
---
fs/lustre/osc/osc_cache.c | 1 -
1 file changed, 1 deletion(-)
diff --git a/fs/lustre/osc/osc_cache.c b/fs/lustre/osc/osc_cache.c
index 3189eb3..2ed7ca2 100644
--- a/fs/lustre/osc/osc_cache.c
+++ b/fs/lustre/osc/osc_cache.c
@@ -1570,7 +1570,6 @@ static bool osc_enter_cache_try(struct client_obd *cli,
cmd1, cmd2) \
({ \
long __ret = timeout; \
- might_sleep(); \
if (!___wait_cond_timeout(condition)) \
__ret = __wait_event_idle_exclusive_timeout_cmd( \
wq_head, condition, timeout, cmd1, cmd2); \
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:53 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:53 -0500
Subject: [lustre-devel] [PATCH 005/622] lustre: llite: return compatible
fsid for statfs
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-6-git-send-email-jsimmons@infradead.org>
From: Fan Yong
Lustre uses 64-bits inode number to identify object on client side.
When re-export Lustre via NFS, NFS will detect whether support fsid
via statfs(). For the non-support case, it will only recognizes and
packs low 32-bits inode number in nfs handle. Such handle cannot be
used to locate the object properly.
To avoid patch linux kernel, Lustre client should generate fsid and
return it via statfs() to up layer.
To be compatible with old Lustre client (NFS server), the fsid will
be generated from super_block::s_dev.
WC-bug-id: https://jira.whamcloud.com/browse/LU-2904
Lustre-commit: abe4d83fab00 ("LU-2904 llite: return compatible fsid for statfs")
Signed-off-by: Fan Yong
Reviewed-on: http://review.whamcloud.com/7434
Reviewed-by: Bobi Jam
Reviewed-by: Jian Yu
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/llite_internal.h | 3 ---
fs/lustre/llite/llite_lib.c | 8 ++++----
fs/lustre/llite/llite_nfs.c | 16 ----------------
3 files changed, 4 insertions(+), 23 deletions(-)
diff --git a/fs/lustre/llite/llite_internal.h b/fs/lustre/llite/llite_internal.h
index f0a50fc..3192340 100644
--- a/fs/lustre/llite/llite_internal.h
+++ b/fs/lustre/llite/llite_internal.h
@@ -538,8 +538,6 @@ struct ll_sb_info {
/* st_blksize returned by stat(2), when non-zero */
unsigned int ll_stat_blksize;
- __kernel_fsid_t ll_fsid;
-
struct kset ll_kset; /* sysfs object */
struct completion ll_kobj_unregister;
};
@@ -941,7 +939,6 @@ static inline ssize_t ll_lov_user_md_size(const struct lov_user_md *lum)
/* llite/llite_nfs.c */
extern const struct export_operations lustre_export_operations;
u32 get_uuid2int(const char *name, int len);
-void get_uuid2fsid(const char *name, int len, __kernel_fsid_t *fsid);
struct inode *search_inode_for_lustre(struct super_block *sb,
const struct lu_fid *fid);
int ll_dir_get_parent_fid(struct inode *dir, struct lu_fid *parent_fid);
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index a48d753..e1932ae 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -591,10 +591,8 @@ static int client_common_fill_super(struct super_block *sb, char *md, char *dt)
* only a node-local comparison.
*/
uuid = obd_get_uuid(sbi->ll_md_exp);
- if (uuid) {
+ if (uuid)
sb->s_dev = get_uuid2int(uuid->uuid, strlen(uuid->uuid));
- get_uuid2fsid(uuid->uuid, strlen(uuid->uuid), &sbi->ll_fsid);
- }
kfree(data);
kfree(osfs);
@@ -1775,6 +1773,7 @@ int ll_statfs(struct dentry *de, struct kstatfs *sfs)
{
struct super_block *sb = de->d_sb;
struct obd_statfs osfs;
+ u64 fsid = huge_encode_dev(sb->s_dev);
int rc;
CDEBUG(D_VFSTRACE, "VFS Op: at %llu jiffies\n", get_jiffies_64());
@@ -1805,7 +1804,8 @@ int ll_statfs(struct dentry *de, struct kstatfs *sfs)
sfs->f_blocks = osfs.os_blocks;
sfs->f_bfree = osfs.os_bfree;
sfs->f_bavail = osfs.os_bavail;
- sfs->f_fsid = ll_s2sbi(sb)->ll_fsid;
+ sfs->f_fsid.val[0] = (u32)fsid;
+ sfs->f_fsid.val[1] = (u32)(fsid >> 32);
return 0;
}
diff --git a/fs/lustre/llite/llite_nfs.c b/fs/lustre/llite/llite_nfs.c
index d6643d0..434f92b 100644
--- a/fs/lustre/llite/llite_nfs.c
+++ b/fs/lustre/llite/llite_nfs.c
@@ -57,22 +57,6 @@ u32 get_uuid2int(const char *name, int len)
return (key0 << 1);
}
-void get_uuid2fsid(const char *name, int len, __kernel_fsid_t *fsid)
-{
- u64 key = 0, key0 = 0x12a3fe2d, key1 = 0x37abe8f9;
-
- while (len--) {
- key = key1 + (key0 ^ (*name++ * 7152373));
- if (key & 0x8000000000000000ULL)
- key -= 0x7fffffffffffffffULL;
- key1 = key0;
- key0 = key;
- }
-
- fsid->val[0] = key;
- fsid->val[1] = key >> 32;
-}
-
struct inode *search_inode_for_lustre(struct super_block *sb,
const struct lu_fid *fid)
{
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:52 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:52 -0500
Subject: [lustre-devel] [PATCH 004/622] lustre: uapi: sync enum
obd_statfs_state
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-5-git-send-email-jsimmons@infradead.org>
With the drift between the OpenSFS and linux client various
enum obd_statfs_state values where dropped that are transmitted
over the wire. Sync the values.
Signed-off-by: James Simmons
---
include/uapi/linux/lustre/lustre_user.h | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/include/uapi/linux/lustre/lustre_user.h b/include/uapi/linux/lustre/lustre_user.h
index f5474c5..27501a2 100644
--- a/include/uapi/linux/lustre/lustre_user.h
+++ b/include/uapi/linux/lustre/lustre_user.h
@@ -101,9 +101,9 @@
enum obd_statfs_state {
OS_STATE_DEGRADED = 0x00000001, /**< RAID degraded/rebuilding */
OS_STATE_READONLY = 0x00000002, /**< filesystem is read-only */
- OS_STATE_RDONLY_1 = 0x00000004, /**< obsolete 1.6, was EROFS=30 */
- OS_STATE_RDONLY_2 = 0x00000008, /**< obsolete 1.6, was EROFS=30 */
- OS_STATE_RDONLY_3 = 0x00000010, /**< obsolete 1.6, was EROFS=30 */
+ OS_STATE_NOPRECREATE = 0x00000004, /**< no object precreation */
+ OS_STATE_ENOSPC = 0x00000020, /**< not enough free space */
+ OS_STATE_ENOINO = 0x00000040, /**< not enough inodes */
};
struct obd_statfs {
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:51 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:51 -0500
Subject: [lustre-devel] [PATCH 003/622] lustre: uapi: remove enum
hsm_progress_states
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-4-git-send-email-jsimmons@infradead.org>
This enum is used only by server side code.
Signed-off-by: James Simmons
---
include/uapi/linux/lustre/lustre_user.h | 21 ---------------------
1 file changed, 21 deletions(-)
diff --git a/include/uapi/linux/lustre/lustre_user.h b/include/uapi/linux/lustre/lustre_user.h
index 0566afad..f5474c5 100644
--- a/include/uapi/linux/lustre/lustre_user.h
+++ b/include/uapi/linux/lustre/lustre_user.h
@@ -1532,27 +1532,6 @@ enum hsm_states {
*/
#define HSM_FLAGS_MASK (HSM_USER_MASK | HSM_STATUS_MASK)
-/**
- * HSM��request progress state
- */
-enum hsm_progress_states {
- HPS_WAITING = 1,
- HPS_RUNNING = 2,
- HPS_DONE = 3,
-};
-
-#define HPS_NONE 0
-
-static inline const char *hsm_progress_state2name(enum hsm_progress_states s)
-{
- switch (s) {
- case HPS_WAITING: return "waiting";
- case HPS_RUNNING: return "running";
- case HPS_DONE: return "done";
- default: return "unknown";
- }
-}
-
struct hsm_extent {
__u64 offset;
__u64 length;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:54 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:54 -0500
Subject: [lustre-devel] [PATCH 006/622] lustre: ldlm: Make kvzalloc | kvfree
use consistent
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-7-git-send-email-jsimmons@infradead.org>
From: "Christopher J. Morrone"
struct ldlm_lock's l_lvb_data field is freed in ldlm_lock_put()
using kfree. However, some other code paths can attach
a buffer to l_lvb_data that was allocated using vmalloc().
This can lead to a kfree() of a vmalloc()ed buffer, which can
trigger a kernel Oops.
WC-bug-id: https://jira.whamcloud.com/browse/LU-4194
Lustre-commit: 9c4d506c5fea ("LU-4194 ldlm: Make OBD_[ALLOC|FREE]_LARGE use consistent")
Signed-off-by: Christopher J. Morrone
Reviewed-on: http://review.whamcloud.com/8298
Reviewed-by: Andreas Dilger
Reviewed-by: Faccini Bruno
Signed-off-by: James Simmons
---
fs/lustre/ldlm/ldlm_lock.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/fs/lustre/ldlm/ldlm_lock.c b/fs/lustre/ldlm/ldlm_lock.c
index 6eebf5f..7242cd1 100644
--- a/fs/lustre/ldlm/ldlm_lock.c
+++ b/fs/lustre/ldlm/ldlm_lock.c
@@ -185,7 +185,7 @@ void ldlm_lock_put(struct ldlm_lock *lock)
lock->l_export = NULL;
}
- kfree(lock->l_lvb_data);
+ kvfree(lock->l_lvb_data);
lu_ref_fini(&lock->l_reference);
OBD_FREE_RCU(lock, sizeof(*lock), &lock->l_handle);
@@ -1548,7 +1548,7 @@ struct ldlm_lock *ldlm_lock_create(struct ldlm_namespace *ns,
if (lvb_len) {
lock->l_lvb_len = lvb_len;
- lock->l_lvb_data = kzalloc(lvb_len, GFP_NOFS);
+ lock->l_lvb_data = kvzalloc(lvb_len, GFP_NOFS);
if (!lock->l_lvb_data) {
rc = -ENOMEM;
goto out;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:55 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:55 -0500
Subject: [lustre-devel] [PATCH 007/622] lustre: llite: limit smallest
max_cached_mb value
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-8-git-send-email-jsimmons@infradead.org>
From: James Nunez
Currently, ost-survey hangs due to calling
'lfs setstripe' in an old (positional) style and
setting max_cached_mb to zero.
In ll_max_cached_mb_seq_write(), the number of
pages requested is set to the max of pages requested
or PTLRPC_MAX_BRW_PAGES to allow the client to make
well formed RPCs.
WC-bug-id: https://jira.whamcloud.com/browse/LU-4768
Lustre-commit: 46bec835ac72 ("LU-4768 tests: Update ost-survey script")
Signed-off-by: James Nunez
Reviewed-on: http://review.whamcloud.com/11971
Reviewed-by: Nathaniel Clark
Reviewed-by: Cliff White
Reviewed-by: Jian Yu
Reviewed-by: Jinshan Xiong
Reviewed-by: James Simmons
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/lproc_llite.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/fs/lustre/llite/lproc_llite.c b/fs/lustre/llite/lproc_llite.c
index e108326..5ac6689 100644
--- a/fs/lustre/llite/lproc_llite.c
+++ b/fs/lustre/llite/lproc_llite.c
@@ -527,6 +527,8 @@ static ssize_t ll_max_cached_mb_seq_write(struct file *file,
totalram_pages() >> (20 - PAGE_SHIFT));
return -ERANGE;
}
+ /* Allow enough cache so clients can make well-formed RPCs */
+ pages_number = max_t(long, pages_number, PTLRPC_MAX_BRW_PAGES);
spin_lock(&sbi->ll_lock);
diff = pages_number - cache->ccc_lru_max;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:01 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:01 -0500
Subject: [lustre-devel] [PATCH 013/622] lustre: obdclass: fix
llog_cat_cleanup() usage on Client
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-14-git-send-email-jsimmons@infradead.org>
From: Bruno Faccini
With patch/commit 3a83b4b9 for LU-5195, LLOG code has been
strengthen against catalog inconsistency by detecting a
referenced plain LLOG is missing and by clearing its
associated entry by calling llog_cat_cleanup(), which now
needs to handle the case where it is also executed on a Client
(ie, cathandle->lgh_obj == NULL) and thus must not attempt to
update on-disk catalog.
WC-bug-id: https://jira.whamcloud.com/browse/LU-6471
Lustre-commit: 485f3ba87433 ("LU-6471 obdclass: fix llog_cat_cleanup() usage on Client")
Signed-off-by: Bruno Faccini
Reviewed-on: http://review.whamcloud.com/14489
Reviewed-by: Alex Zhuravlev
Reviewed-by: John L. Hammond
Reviewed-by: Mikhail Pershin
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/obdclass/llog_cat.c | 6 ++----
1 file changed, 2 insertions(+), 4 deletions(-)
diff --git a/fs/lustre/obdclass/llog_cat.c b/fs/lustre/obdclass/llog_cat.c
index 580d807..ca97e08 100644
--- a/fs/lustre/obdclass/llog_cat.c
+++ b/fs/lustre/obdclass/llog_cat.c
@@ -133,10 +133,8 @@ int llog_cat_close(const struct lu_env *env, struct llog_handle *cathandle)
list_del_init(&loghandle->u.phd.phd_entry);
llog_close(env, loghandle);
}
- /* if handle was stored in ctxt, remove it too */
- if (cathandle->lgh_ctxt->loc_handle == cathandle)
- cathandle->lgh_ctxt->loc_handle = NULL;
- return llog_close(env, cathandle);
+
+ return 0;
}
EXPORT_SYMBOL(llog_cat_close);
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:05 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:05 -0500
Subject: [lustre-devel] [PATCH 017/622] lustre: obdclass: new wrapper to
convert NID to string
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-18-git-send-email-jsimmons@infradead.org>
From: Liang Zhen
This patch includes a couple of changes:
- add new wrapper function obd_import_nid2str
- use obd_import_nid2str and obd_export_nid2str to replace all
libcfs_nid2str conversions for NID of export/import connection
WC-bug-id: https://jira.whamcloud.com/browse/LU-6032
Lustre-commit: 61f9847a812f ("LU-6032 obdclass: new wrapper to convert NID to string")
Signed-off-by: Liang Zhen
Reviewed-on: https://review.whamcloud.com/12956
Reviewed-by: Dmitry Eremin
Reviewed-by: Amir Shehata
Reviewed-by: James Simmons
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/include/obd_class.h | 12 ++++++++++++
fs/lustre/ldlm/ldlm_lock.c | 4 ++--
fs/lustre/ptlrpc/client.c | 5 ++---
fs/lustre/ptlrpc/import.c | 6 +++---
4 files changed, 19 insertions(+), 8 deletions(-)
diff --git a/fs/lustre/include/obd_class.h b/fs/lustre/include/obd_class.h
index 146c37e..d896049 100644
--- a/fs/lustre/include/obd_class.h
+++ b/fs/lustre/include/obd_class.h
@@ -86,6 +86,18 @@ struct obd_device *class_devices_in_group(struct obd_uuid *grp_uuid,
int obd_connect_flags2str(char *page, int count, u64 flags, u64 flags2,
const char *sep);
+static inline char *obd_export_nid2str(struct obd_export *exp)
+{
+ return exp->exp_connection ?
+ libcfs_nid2str(exp->exp_connection->c_peer.nid) : "";
+}
+
+static inline char *obd_import_nid2str(struct obd_import *imp)
+{
+ return imp->imp_connection ?
+ libcfs_nid2str(imp->imp_connection->c_peer.nid) : "";
+}
+
int obd_zombie_impexp_init(void);
void obd_zombie_impexp_stop(void);
void obd_zombie_barrier(void);
diff --git a/fs/lustre/ldlm/ldlm_lock.c b/fs/lustre/ldlm/ldlm_lock.c
index 7242cd1..aa19b89 100644
--- a/fs/lustre/ldlm/ldlm_lock.c
+++ b/fs/lustre/ldlm/ldlm_lock.c
@@ -1987,11 +1987,11 @@ void _ldlm_lock_debug(struct ldlm_lock *lock,
vaf.va = &args;
if (exp && exp->exp_connection) {
- nid = libcfs_nid2str(exp->exp_connection->c_peer.nid);
+ nid = obd_export_nid2str(exp);
} else if (exp && exp->exp_obd) {
struct obd_import *imp = exp->exp_obd->u.cli.cl_import;
- nid = libcfs_nid2str(imp->imp_connection->c_peer.nid);
+ nid = obd_import_nid2str(imp);
}
if (!resource) {
diff --git a/fs/lustre/ptlrpc/client.c b/fs/lustre/ptlrpc/client.c
index a533cbb..424db55 100644
--- a/fs/lustre/ptlrpc/client.c
+++ b/fs/lustre/ptlrpc/client.c
@@ -1605,8 +1605,7 @@ static int ptlrpc_send_new_req(struct ptlrpc_request *req)
current->comm,
imp->imp_obd->obd_uuid.uuid,
lustre_msg_get_status(req->rq_reqmsg), req->rq_xid,
- libcfs_nid2str(imp->imp_connection->c_peer.nid),
- lustre_msg_get_opc(req->rq_reqmsg));
+ obd_import_nid2str(imp), lustre_msg_get_opc(req->rq_reqmsg));
rc = ptl_send_rpc(req, 0);
if (rc == -ENOMEM) {
@@ -2017,7 +2016,7 @@ int ptlrpc_check_set(const struct lu_env *env, struct ptlrpc_request_set *set)
current->comm, imp->imp_obd->obd_uuid.uuid,
lustre_msg_get_status(req->rq_reqmsg),
req->rq_xid,
- libcfs_nid2str(imp->imp_connection->c_peer.nid),
+ obd_import_nid2str(imp),
lustre_msg_get_opc(req->rq_reqmsg));
spin_lock(&imp->imp_lock);
diff --git a/fs/lustre/ptlrpc/import.c b/fs/lustre/ptlrpc/import.c
index d032962..dca4aa0 100644
--- a/fs/lustre/ptlrpc/import.c
+++ b/fs/lustre/ptlrpc/import.c
@@ -171,13 +171,13 @@ int ptlrpc_set_import_discon(struct obd_import *imp, u32 conn_cnt)
LCONSOLE_WARN("%s: Connection to %.*s (at %s) was lost; in progress operations using this service will wait for recovery to complete\n",
imp->imp_obd->obd_name,
target_len, target_start,
- libcfs_nid2str(imp->imp_connection->c_peer.nid));
+ obd_import_nid2str(imp));
} else {
LCONSOLE_ERROR_MSG(0x166,
"%s: Connection to %.*s (at %s) was lost; in progress operations using this service will fail\n",
imp->imp_obd->obd_name,
target_len, target_start,
- libcfs_nid2str(imp->imp_connection->c_peer.nid));
+ obd_import_nid2str(imp));
}
IMPORT_SET_STATE_NOLOCK(imp, LUSTRE_IMP_DISCON);
spin_unlock(&imp->imp_lock);
@@ -1461,7 +1461,7 @@ int ptlrpc_import_recovery_state_machine(struct obd_import *imp)
LCONSOLE_INFO("%s: Connection restored to %.*s (at %s)\n",
imp->imp_obd->obd_name,
target_len, target_start,
- libcfs_nid2str(imp->imp_connection->c_peer.nid));
+ obd_import_nid2str(imp));
}
if (imp->imp_state == LUSTRE_IMP_FULL) {
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:08 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:08 -0500
Subject: [lustre-devel] [PATCH 020/622] lnet: libcfs: remove unnecessary
set_fs(KERNEL_DS)
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-21-git-send-email-jsimmons@infradead.org>
From: Mike Marciniszyn
When we converted to using kernel_write(), we left some
set_fs() calls that are not unnecessary.
Remove them.
Original OpenSFS version of this patch, as mentioned below,
did the full conversion to kernel_write.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10560
lustre-commit: b9a32054600a ("LU-10560 libcfs: Use kernel_write when appropriate")
Signed-off-by: Mike Marciniszyn
Reviewed-on: https://review.whamcloud.com/31154
Reviewed-by: James Simmons
Reviewed-by: Dmitry Eremin
Reviewed-by: John L. Hammond
Reviewed-by: Oleg Drokin
igned-off-by: James Simmons
---
net/lnet/libcfs/tracefile.c | 5 +----
1 file changed, 1 insertion(+), 4 deletions(-)
diff --git a/net/lnet/libcfs/tracefile.c b/net/lnet/libcfs/tracefile.c
index 3b29116..6e4cc31 100644
--- a/net/lnet/libcfs/tracefile.c
+++ b/net/lnet/libcfs/tracefile.c
@@ -807,7 +807,6 @@ int cfs_tracefile_dump_all_pages(char *filename)
struct cfs_trace_page *tage;
struct cfs_trace_page *tmp;
char *buf;
- mm_segment_t __oldfs;
int rc;
down_write(&cfs_tracefile_sem);
@@ -828,8 +827,6 @@ int cfs_tracefile_dump_all_pages(char *filename)
rc = 0;
goto close;
}
- __oldfs = get_fs();
- set_fs(KERNEL_DS);
/* ok, for now, just write the pages. in the future we'll be building
* iobufs with the pages and calling generic_direct_IO
@@ -851,7 +848,7 @@ int cfs_tracefile_dump_all_pages(char *filename)
list_del(&tage->linkage);
cfs_tage_free(tage);
}
- set_fs(__oldfs);
+
rc = vfs_fsync(filp, 1);
if (rc)
pr_err("sync returns %d\n", rc);
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:06 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:06 -0500
Subject: [lustre-devel] [PATCH 018/622] lustre: ptlrpc: Add QoS for uid and
gid in NRS-TBF
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-19-git-send-email-jsimmons@infradead.org>
From: Teddy Chan
This patch add a new QoS feature in TBF policy which could
limits the rate based on uid or gid. The policy is able to
limit the rate both on MDT and OSS site.
The command for this feature is like:
Start the tbf uid QoS on OST:
lctl set_param ost.OSS.*.nrs_policies="tbf uid"
Limit the rate of ptlrpc requests of the uid 500
lctl set_param ost.OSS.*.nrs_tbf_rule=
"start tbf_name uid={500} rate=100"
Start the tbf gid QoS on OST:
lctl set_param ost.OSS.*.nrs_policies="tbf gid"
Limit the rate of ptlrpc requests of the gid 500
lctl set_param ost.OSS.*.nrs_tbf_rule=
"start tbf_name gid={500} rate=100"
or use generic tbf rule to mix them on OST:
lctl set_param ost.OSS.*.nrs_policies="tbf"
Limit the rate of ptlrpc requests of the uid 500 gid 500
lctl set_param ost.OSS.*.nrs_tbf_rule=
"start tbf_name uid={500}&gid={500} rate=100"
Also, you can use the following rule to control all reqs
to mds:
Start the tbf uid QoS on MDS:
lctl set_param mds.MDS.*.nrs_policies="tbf uid"
Limit the rate of ptlrpc requests of the uid 500
lctl set_param mds.MDS.*.nrs_tbf_rule=
"start tbf_name uid={500} rate=100"
For the linux client we need to send the uid and gid
information to the NRS-TBF handling on the servers.
WC-bug-id: https://jira.whamcloud.com/browse/LU-9658
Lustre-commit: e0cdde123c14 ("LU-9658 ptlrpc: Add QoS for uid and gid in NRS-TBF")
Signed-off-by: Teddy Chan
Signed-off-by: Li Xi
Signed-off-by: Wang Shilong
Signed-off-by: Qian Yingjin
Reviewed-on: https://review.whamcloud.com/27608
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/vvp_object.c | 5 ++---
fs/lustre/obdclass/obdo.c | 5 +++++
fs/lustre/osc/osc_request.c | 10 ++++++++++
3 files changed, 17 insertions(+), 3 deletions(-)
diff --git a/fs/lustre/llite/vvp_object.c b/fs/lustre/llite/vvp_object.c
index 24cde0d..eeb8823 100644
--- a/fs/lustre/llite/vvp_object.c
+++ b/fs/lustre/llite/vvp_object.c
@@ -196,7 +196,7 @@ static int vvp_object_glimpse(const struct lu_env *env,
static void vvp_req_attr_set(const struct lu_env *env, struct cl_object *obj,
struct cl_req_attr *attr)
{
- u64 valid_flags = OBD_MD_FLTYPE;
+ u64 valid_flags = OBD_MD_FLTYPE | OBD_MD_FLUID | OBD_MD_FLGID;
struct inode *inode;
struct obdo *oa;
@@ -204,8 +204,7 @@ static void vvp_req_attr_set(const struct lu_env *env, struct cl_object *obj,
inode = vvp_object_inode(obj);
if (attr->cra_type == CRT_WRITE) {
- valid_flags |= OBD_MD_FLMTIME | OBD_MD_FLCTIME |
- OBD_MD_FLUID | OBD_MD_FLGID;
+ valid_flags |= OBD_MD_FLMTIME | OBD_MD_FLCTIME;
obdo_set_o_projid(oa, ll_i2info(inode)->lli_projid);
}
obdo_from_inode(oa, inode, valid_flags & attr->cra_flags);
diff --git a/fs/lustre/obdclass/obdo.c b/fs/lustre/obdclass/obdo.c
index 1926896..e5475f1 100644
--- a/fs/lustre/obdclass/obdo.c
+++ b/fs/lustre/obdclass/obdo.c
@@ -144,6 +144,11 @@ void lustre_set_wire_obdo(const struct obd_connect_data *ocd,
if (!ocd)
return;
+ if (!(wobdo->o_valid & OBD_MD_FLUID))
+ wobdo->o_uid = from_kuid(&init_user_ns, current_uid());
+ if (!(wobdo->o_valid & OBD_MD_FLGID))
+ wobdo->o_gid = from_kgid(&init_user_ns, current_gid());
+
if (unlikely(!(ocd->ocd_connect_flags & OBD_CONNECT_FID)) &&
fid_seq_is_echo(ostid_seq(&lobdo->o_oi))) {
/*
diff --git a/fs/lustre/osc/osc_request.c b/fs/lustre/osc/osc_request.c
index 300dee5..99c9620 100644
--- a/fs/lustre/osc/osc_request.c
+++ b/fs/lustre/osc/osc_request.c
@@ -1184,6 +1184,16 @@ static int osc_brw_prep_request(int cmd, struct client_obd *cli,
lustre_set_wire_obdo(&req->rq_import->imp_connect_data, &body->oa, oa);
+ /* For READ and WRITE, we can't fill o_uid and o_gid using from_kuid()
+ * and from_kgid(), because they are asynchronous. Fortunately, variable
+ * oa contains valid o_uid and o_gid in these two operations.
+ * Besides, filling o_uid and o_gid is enough for nrs-tbf, see LU-9658.
+ * OBD_MD_FLUID and OBD_MD_FLUID is not set in order to avoid breaking
+ * other process logic
+ */
+ body->oa.o_uid = oa->o_uid;
+ body->oa.o_gid = oa->o_gid;
+
obdo_to_ioobj(oa, ioobj);
ioobj->ioo_bufcnt = niocount;
/* The high bits of ioo_max_brw tells server _maximum_ number of bulks
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:10 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:10 -0500
Subject: [lustre-devel] [PATCH 022/622] lustre: llite: yield cpu after call
to ll_agl_trigger
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-23-git-send-email-jsimmons@infradead.org>
From: Ann Koehler
The statahead and agl threads loop over all entries in the
directory without yielding the CPU. If the number of entries in
the directory is large enough then these threads may trigger
soft lockups. The fix is to add calls to cond_resched() after
calling ll_agl_trigger(), which gets the glimpse lock for a
file.
Cray-bug-id: LUS-2584
WC-bug-id: https://jira.whamcloud.com/browse/LU-10649
Lustre-commit: 031001f0d438 ("LU-10649 llite: yield cpu after call to ll_agl_trigger")
Signed-off-by: Ann Koehler
Signed-off-by: Chris Horn
Reviewed-on: https://review.whamcloud.com/31240
Reviewed-by: Patrick Farrell
Reviewed-by: Sergey Cheremencev
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/statahead.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/fs/lustre/llite/statahead.c b/fs/lustre/llite/statahead.c
index 99b3fee..4a61dac 100644
--- a/fs/lustre/llite/statahead.c
+++ b/fs/lustre/llite/statahead.c
@@ -907,6 +907,7 @@ static int ll_agl_thread(void *arg)
list_del_init(&clli->lli_agl_list);
spin_unlock(&plli->lli_agl_lock);
ll_agl_trigger(&clli->lli_vfs_inode, sai);
+ cond_resched();
} else {
spin_unlock(&plli->lli_agl_lock);
}
@@ -1071,7 +1072,7 @@ static int ll_statahead_thread(void *arg)
ll_agl_trigger(&clli->lli_vfs_inode,
sai);
-
+ cond_resched();
spin_lock(&lli->lli_agl_lock);
}
spin_unlock(&lli->lli_agl_lock);
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:15 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:15 -0500
Subject: [lustre-devel] [PATCH 027/622] lustre: lu_object: improve debug
message for lu_object_put()
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-28-git-send-email-jsimmons@infradead.org>
From: Alexey Lyashkov
Use a top level object in debug in lu_object_put to match with
lu_object_get.
WC-bug-id: https://jira.whamcloud.com/browse/LU-LU-10877
Lustre-commit: fd669eba1921 ("LU-10877 lu: fix reference leak")
Signed-off-by: Alexey Lyashkov
Reviewed-on: https://review.whamcloud.com/31870
Reviewed-by: Andrew Perepechko
Reviewed-by: Sergey Cheremencev
Reviewed-by: Alex Zhuravlev
Reviewed-by: Mikhal Pershin
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/obdclass/lu_object.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/fs/lustre/obdclass/lu_object.c b/fs/lustre/obdclass/lu_object.c
index d8dfc721..2ab4977 100644
--- a/fs/lustre/obdclass/lu_object.c
+++ b/fs/lustre/obdclass/lu_object.c
@@ -184,8 +184,8 @@ void lu_object_put(const struct lu_env *env, struct lu_object *o)
LASSERT(list_empty(&top->loh_lru));
list_add_tail(&top->loh_lru, &bkt->lsb_lru);
percpu_counter_inc(&site->ls_lru_len_counter);
- CDEBUG(D_INODE, "Add %p to site lru. hash: %p, bkt: %p\n",
- o, site->ls_obj_hash, bkt);
+ CDEBUG(D_INODE, "Add %p/%p to site lru. hash: %p, bkt: %p\n",
+ orig, top, site->ls_obj_hash, bkt);
cfs_hash_bd_unlock(site->ls_obj_hash, &bd, 1);
return;
}
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:16 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:16 -0500
Subject: [lustre-devel] [PATCH 028/622] lustre: idl: remove obsolete
directory split flags
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-29-git-send-email-jsimmons@infradead.org>
From: Andreas Dilger
The directory split functionality from the old CMD (pre-DNE)
feature was never usable in production, and was removed before
the DNE 2.4 release. Remove old flags relating to this feature.
WC-bug-id: https://jira.whamcloud.com/browse/LU-1187
Lustre-commit: 5c53c353fd82 ("LU-1187 idl: remove obsolete directory split flags")
Signed-off-by: Andreas Dilger
Reviewed-on: https://review.whamcloud.com/31700
Reviewed-by: James Simmons
Reviewed-by: Lai Siyao
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/mdc/mdc_lib.c | 2 --
fs/lustre/ptlrpc/wiretest.c | 4 ----
include/uapi/linux/lustre/lustre_idl.h | 4 ++--
3 files changed, 2 insertions(+), 8 deletions(-)
diff --git a/fs/lustre/mdc/mdc_lib.c b/fs/lustre/mdc/mdc_lib.c
index d4b2bb9..467503c 100644
--- a/fs/lustre/mdc/mdc_lib.c
+++ b/fs/lustre/mdc/mdc_lib.c
@@ -520,8 +520,6 @@ void mdc_getattr_pack(struct ptlrpc_request *req, u64 valid, u32 flags,
&RMF_MDT_BODY);
b->mbo_valid = valid;
- if (op_data->op_bias & MDS_CHECK_SPLIT)
- b->mbo_valid |= OBD_MD_FLCKSPLIT;
if (op_data->op_bias & MDS_CROSS_REF)
b->mbo_valid |= OBD_MD_FLCROSSREF;
b->mbo_eadatasize = ea_size;
diff --git a/fs/lustre/ptlrpc/wiretest.c b/fs/lustre/ptlrpc/wiretest.c
index 21698cc..bcd0229 100644
--- a/fs/lustre/ptlrpc/wiretest.c
+++ b/fs/lustre/ptlrpc/wiretest.c
@@ -1341,8 +1341,6 @@ void lustre_assert_wire_constants(void)
OBD_MD_FLMDSCAPA);
LASSERTF(OBD_MD_FLOSSCAPA == (0x0000040000000000ULL), "found 0x%.16llxULL\n",
OBD_MD_FLOSSCAPA);
- LASSERTF(OBD_MD_FLCKSPLIT == (0x0000080000000000ULL), "found 0x%.16llxULL\n",
- OBD_MD_FLCKSPLIT);
LASSERTF(OBD_MD_FLCROSSREF == (0x0000100000000000ULL), "found 0x%.16llxULL\n",
OBD_MD_FLCROSSREF);
LASSERTF(OBD_MD_FLGETATTRLOCK == (0x0000200000000000ULL), "found 0x%.16llxULL\n",
@@ -1866,8 +1864,6 @@ void lustre_assert_wire_constants(void)
LASSERTF((int)sizeof(((struct ll_fid *)0)->f_type) == 4, "found %lld\n",
(long long)(int)sizeof(((struct ll_fid *)0)->f_type));
- LASSERTF(MDS_CHECK_SPLIT == 0x00000001UL, "found 0x%.8xUL\n",
- (unsigned int)MDS_CHECK_SPLIT);
LASSERTF(MDS_CROSS_REF == 0x00000002UL, "found 0x%.8xUL\n",
(unsigned int)MDS_CROSS_REF);
LASSERTF(MDS_VTX_BYPASS == 0x00000004UL, "found 0x%.8xUL\n",
diff --git a/include/uapi/linux/lustre/lustre_idl.h b/include/uapi/linux/lustre/lustre_idl.h
index 0bce63d..589bb81 100644
--- a/include/uapi/linux/lustre/lustre_idl.h
+++ b/include/uapi/linux/lustre/lustre_idl.h
@@ -1131,7 +1131,7 @@ static inline __u32 lov_mds_md_size(__u16 stripes, __u32 lmm_magic)
/* OBD_MD_FLRMTPERM (0x0000010000000000ULL) remote perm, obsolete */
#define OBD_MD_FLMDSCAPA (0x0000020000000000ULL) /* MDS capability */
#define OBD_MD_FLOSSCAPA (0x0000040000000000ULL) /* OSS capability */
-#define OBD_MD_FLCKSPLIT (0x0000080000000000ULL) /* Check split on server */
+/* OBD_MD_FLCKSPLIT (0x0000080000000000ULL) obsolete 2.3.58*/
#define OBD_MD_FLCROSSREF (0x0000100000000000ULL) /* Cross-ref case */
#define OBD_MD_FLGETATTRLOCK (0x0000200000000000ULL) /* Get IOEpoch attributes
* under lock; for xattr
@@ -1640,7 +1640,7 @@ struct mdt_rec_setattr {
#define MDS_ATTR_PROJID 0x10000ULL /* = 65536 */
enum mds_op_bias {
- MDS_CHECK_SPLIT = 1 << 0,
+/* MDS_CHECK_SPLIT = 1 << 0, obsolete before 2.3.58 */
MDS_CROSS_REF = 1 << 1,
MDS_VTX_BYPASS = 1 << 2,
MDS_PERM_BYPASS = 1 << 3,
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:19 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:19 -0500
Subject: [lustre-devel] [PATCH 031/622] lustre: ldlm: change
LDLM_POOL_ADD_VAR macro to inline function
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-32-git-send-email-jsimmons@infradead.org>
Simple cleanup to create inline funciton ldlm_pool_add_var().
WC-bug-id: https://jira.hpdd.intel.com/browse/LU-8066
Lustre-commit: 05a36534ba2d ("LU-8066 ldlm: move all remaining files from procfs to debugfs")
Signed-off-by: Dmitry Eremin
Signed-off-by: Oleg Drokin
Signed-off-by: James Simmons
Reviewed-on: https://review.whamcloud.com/29255
WC-bug-id: https://jira.hpdd.intel.com/browse/LU-3319
Lustre-commit: 4ad445ccd54 ("LU-3319 procfs: move ldlm proc handling over to seq_file")
Reviewed-on: http://review.whamcloud.com/7293
Reviewed-by: Dmitry Eremin
Reviewed-by: Andreas Dilger
Reviewed-by: Peng Tao
Reviewed-by: Bob Glossman
Reviewed-by: Yang Sheng
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/ldlm/ldlm_internal.h | 10 ++++++++++
fs/lustre/ldlm/ldlm_pool.c | 11 ++---------
2 files changed, 12 insertions(+), 9 deletions(-)
diff --git a/fs/lustre/ldlm/ldlm_internal.h b/fs/lustre/ldlm/ldlm_internal.h
index 6e54521..96dff1d 100644
--- a/fs/lustre/ldlm/ldlm_internal.h
+++ b/fs/lustre/ldlm/ldlm_internal.h
@@ -292,6 +292,16 @@ enum ldlm_policy_res {
} \
struct __##var##__dummy_write {; } /* semicolon catcher */
+static inline void
+ldlm_add_var(struct lprocfs_vars *vars, struct dentry *debugfs_entry,
+ const char *name, void *data, const struct file_operations *ops)
+{
+ vars->name = name;
+ vars->data = data;
+ vars->fops = ops;
+ ldebugfs_add_vars(debugfs_entry, vars, NULL);
+}
+
static inline int is_granted_or_cancelled(struct ldlm_lock *lock)
{
int ret = 0;
diff --git a/fs/lustre/ldlm/ldlm_pool.c b/fs/lustre/ldlm/ldlm_pool.c
index 04bf5de..d2149a6 100644
--- a/fs/lustre/ldlm/ldlm_pool.c
+++ b/fs/lustre/ldlm/ldlm_pool.c
@@ -504,14 +504,6 @@ static ssize_t grant_speed_show(struct kobject *kobj, struct attribute *attr,
LDLM_POOL_SYSFS_WRITER_NOLOCK_STORE(lock_volume_factor, atomic);
LUSTRE_RW_ATTR(lock_volume_factor);
-#define LDLM_POOL_ADD_VAR(_name, var, ops) \
- do { \
- pool_vars[0].name = #_name; \
- pool_vars[0].data = var; \
- pool_vars[0].fops = ops; \
- ldebugfs_add_vars(pl->pl_debugfs_entry, pool_vars, NULL);\
- } while (0)
-
/* These are for pools in /sys/fs/lustre/ldlm/namespaces/.../pool */
static struct attribute *ldlm_pl_attrs[] = {
&lustre_attr_grant_speed.attr,
@@ -571,7 +563,8 @@ static int ldlm_pool_debugfs_init(struct ldlm_pool *pl)
memset(pool_vars, 0, sizeof(pool_vars));
- LDLM_POOL_ADD_VAR(state, pl, &lprocfs_pool_state_fops);
+ ldlm_add_var(&pool_vars[0], pl->pl_debugfs_entry, "state", pl,
+ &lprocfs_pool_state_fops);
pl->pl_stats = lprocfs_alloc_stats(LDLM_POOL_LAST_STAT -
LDLM_POOL_FIRST_STAT, 0);
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:00 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:00 -0500
Subject: [lustre-devel] [PATCH 012/622] lustre: lov: protected ost pool
count updation
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-13-git-send-email-jsimmons@infradead.org>
From: Jadhav Vikram
ASSERTION(iter->lpi_idx <= ((iter->lpi_pool)->pool_obds.op_count)
caused due to reading of ost pool count is not protected in
pool_proc_next and pool_proc_show, pool_proc_show get called when
op_count was zero.
Fix to protect ost pool count by taking lock at start sequence
function pool_proc_start and released lock in pool_proc_stop.
Rather than using down_read / up_read pairs around pool_proc_next
and pool_proc_show, this changes make sure ost pool data gets
protected throughout sequence operation.
Seagate-bug-id: MRP-3629
WC-bug-id: https://jira.whamcloud.com/browse/LU-9620
Lustre-commit: 61c803319b91 ("LU-9620 lod: protected ost pool count updation")
Signed-off-by: Jadhav Vikram
Reviewed-by: Ashish Purkar
Reviewed-by: Vladimir Saveliev
Reviewed-on: https://review.whamcloud.com/27506
Reviewed-by: Fan Yong
Reviewed-by: Niu Yawei
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/lov/lov_pool.c | 9 +++------
1 file changed, 3 insertions(+), 6 deletions(-)
diff --git a/fs/lustre/lov/lov_pool.c b/fs/lustre/lov/lov_pool.c
index 60565b9..a0552fb 100644
--- a/fs/lustre/lov/lov_pool.c
+++ b/fs/lustre/lov/lov_pool.c
@@ -117,14 +117,11 @@ static void *pool_proc_next(struct seq_file *s, void *v, loff_t *pos)
/* iterate to find a non empty entry */
prev_idx = iter->idx;
- down_read(&pool_tgt_rw_sem(iter->pool));
iter->idx++;
- if (iter->idx == pool_tgt_count(iter->pool)) {
+ if (iter->idx >= pool_tgt_count(iter->pool)) {
iter->idx = prev_idx; /* we stay on the last entry */
- up_read(&pool_tgt_rw_sem(iter->pool));
return NULL;
}
- up_read(&pool_tgt_rw_sem(iter->pool));
(*pos)++;
/* return != NULL to continue */
return iter;
@@ -157,6 +154,7 @@ static void *pool_proc_start(struct seq_file *s, loff_t *pos)
*/
/* /!\ do not forget to restore it to pool before freeing it */
s->private = iter;
+ down_read(&pool_tgt_rw_sem(pool));
if (*pos > 0) {
loff_t i;
void *ptr;
@@ -179,6 +177,7 @@ static void pool_proc_stop(struct seq_file *s, void *v)
* we have to free only if s->private is an iterator
*/
if ((iter) && (iter->magic == POOL_IT_MAGIC)) {
+ up_read(&pool_tgt_rw_sem(iter->pool));
/* we restore s->private so next call to pool_proc_start()
* will work
*/
@@ -197,9 +196,7 @@ static int pool_proc_show(struct seq_file *s, void *v)
LASSERT(iter->pool);
LASSERT(iter->idx <= pool_tgt_count(iter->pool));
- down_read(&pool_tgt_rw_sem(iter->pool));
tgt = pool_tgt(iter->pool, iter->idx);
- up_read(&pool_tgt_rw_sem(iter->pool));
if (tgt)
seq_printf(s, "%s\n", obd_uuid2str(&tgt->ltd_uuid));
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:03 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:03 -0500
Subject: [lustre-devel] [PATCH 015/622] lustre: obdclass: allow specifying
complex jobids
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-16-git-send-email-jsimmons@infradead.org>
From: Andreas Dilger
Allow specifying a format string for the jobid_name variable to create
a jobid for processes on the client. The jobid_name is used when
jobid_var=nodelocal, if jobid_name contains "%j", or as a fallback if
getting the specified jobid_var from the environment fails.
The jobid_node string allows the following escape sequences:
%e = executable name
%g = group ID
%h = hostname (system utsname)
%j = jobid from jobid_var environment variable
%p = process ID
%u = user ID
Any unknown escape sequences are dropped. Other arbitrary characters
pass through unmodified, up to the maximum jobid string size of 32,
though whitespace within the jobid is not copied.
This allows, for example, specifying an arbitrary prefix, such as the
cluster name, in addition to the traditional "procname.uid" format,
to distinguish between jobs running on clients in different clusters:
lctl set_param jobid_var=nodelocal jobid_name=cluster2.%e.%u
or
lctl set_param jobid_var=SLURM_JOB_ID jobid_name=cluster2.%j.%e
To use an environment-specified JobID, if available, but fall back to
a static string for all processes that do not have a valid JobID:
lctl set_param jobid_var=SLURM_JOB_ID jobid_name=unknown
Implementation notes:
The LUSTRE_JOBID_SIZE includes a trailing NUL, so don't use
"LUSTRE_JOBID_SIZE + 1" anywhere, as that is misleading.
Rename the "obd_jobid_node" variable to "obd_jobid_name" to match
the sysfs "jobid_name" parameter name to avoid confusion.
Rename "struct jobid_to_pid_map" to "jobid_pid_map" since this is
not actually mapping from a jobid *to* a PID, but the reverse.
Save jobid length, and reorder fields to avoid holes in structure.
Consolidate PID->jobid cache handling in jobid_get_from_cache(),
which only does environment lookups and caches the results.
The fallback to using obd_jobid_name is handled by the caller.
Rename check_job_name() to jobid_name_is_valid(), since that makes
it clear to the reader a "true" return is a valid name.
In jobid_cache_init() there is no benefit for locking the jobid_hash
creation, since the spinlock is just initialized in this function,
so multiple callers of this function would already be broken.
Pass the buffer size from the callers (who know the buffer size) to
lustre_get_jobid() instead of assuming it is LUSTRE_JOBID_SIZE.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10698
Lustre-commit: 6488c0ec57de ("LU-10698 obdclass: allow specifying complex jobids")
Signed-off-by: Andreas Dilger
Reviewed-on: https://review.whamcloud.com/31691
Reviewed-by: Jinshan Xiong
Reviewed-by: Ben Evans
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/include/obd_class.h | 4 +-
fs/lustre/llite/llite_internal.h | 4 +-
fs/lustre/llite/llite_lib.c | 2 +-
fs/lustre/llite/vvp_io.c | 2 +-
fs/lustre/llite/vvp_object.c | 3 +-
fs/lustre/obdclass/jobid.c | 95 +++++++++++++++++++++++++++++++---
fs/lustre/obdclass/obd_sysfs.c | 10 ++--
fs/lustre/ptlrpc/pack_generic.c | 4 +-
include/uapi/linux/lustre/lustre_idl.h | 2 +-
9 files changed, 105 insertions(+), 21 deletions(-)
diff --git a/fs/lustre/include/obd_class.h b/fs/lustre/include/obd_class.h
index 9e07853..146c37e 100644
--- a/fs/lustre/include/obd_class.h
+++ b/fs/lustre/include/obd_class.h
@@ -54,7 +54,7 @@
/* OBD Operations Declarations */
struct obd_device *class_exp2obd(struct obd_export *exp);
int class_handle_ioctl(unsigned int cmd, unsigned long arg);
-int lustre_get_jobid(char *jobid);
+int lustre_get_jobid(char *jobid, size_t len);
struct lu_device_type;
@@ -1672,7 +1672,7 @@ static inline void class_uuid_unparse(class_uuid_t uu, struct obd_uuid *out)
int class_check_uuid(struct obd_uuid *uuid, u64 nid);
/* class_obd.c */
-extern char obd_jobid_node[];
+extern char obd_jobid_name[];
int class_procfs_init(void);
int class_procfs_clean(void);
diff --git a/fs/lustre/llite/llite_internal.h b/fs/lustre/llite/llite_internal.h
index fbe93a4..d0a703d 100644
--- a/fs/lustre/llite/llite_internal.h
+++ b/fs/lustre/llite/llite_internal.h
@@ -195,11 +195,11 @@ struct ll_inode_info {
int lli_async_rc;
/*
- * whenever a process try to read/write the file, the
+ * Whenever a process try to read/write the file, the
* jobid of the process will be saved here, and it'll
* be packed into the write PRC when flush later.
*
- * so the read/write statistics for jobid will not be
+ * So the read/write statistics for jobid will not be
* accurate if the file is shared by different jobs.
*/
char lli_jobid[LUSTRE_JOBID_SIZE];
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index 12aafe0..7580d57 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -937,7 +937,7 @@ void ll_lli_init(struct ll_inode_info *lli)
lli->lli_async_rc = 0;
}
mutex_init(&lli->lli_layout_mutex);
- memset(lli->lli_jobid, 0, LUSTRE_JOBID_SIZE);
+ memset(lli->lli_jobid, 0, sizeof(lli->lli_jobid));
}
int ll_fill_super(struct super_block *sb)
diff --git a/fs/lustre/llite/vvp_io.c b/fs/lustre/llite/vvp_io.c
index 37bf942..85bb3e0 100644
--- a/fs/lustre/llite/vvp_io.c
+++ b/fs/lustre/llite/vvp_io.c
@@ -1419,7 +1419,7 @@ int vvp_io_init(const struct lu_env *env, struct cl_object *obj,
* it's not accurate if the file is shared by different
* jobs.
*/
- lustre_get_jobid(lli->lli_jobid);
+ lustre_get_jobid(lli->lli_jobid, sizeof(lli->lli_jobid));
} else if (io->ci_type == CIT_SETATTR) {
if (!cl_io_is_trunc(io))
io->ci_lockreq = CILR_MANDATORY;
diff --git a/fs/lustre/llite/vvp_object.c b/fs/lustre/llite/vvp_object.c
index c750a80..24cde0d 100644
--- a/fs/lustre/llite/vvp_object.c
+++ b/fs/lustre/llite/vvp_object.c
@@ -212,7 +212,8 @@ static void vvp_req_attr_set(const struct lu_env *env, struct cl_object *obj,
obdo_set_parent_fid(oa, &ll_i2info(inode)->lli_fid);
if (OBD_FAIL_CHECK(OBD_FAIL_LFSCK_INVALID_PFID))
oa->o_parent_oid++;
- memcpy(attr->cra_jobid, ll_i2info(inode)->lli_jobid, LUSTRE_JOBID_SIZE);
+ memcpy(attr->cra_jobid, ll_i2info(inode)->lli_jobid,
+ sizeof(attr->cra_jobid));
}
static const struct cl_object_operations vvp_ops = {
diff --git a/fs/lustre/obdclass/jobid.c b/fs/lustre/obdclass/jobid.c
index 3655a2e..8bad859 100644
--- a/fs/lustre/obdclass/jobid.c
+++ b/fs/lustre/obdclass/jobid.c
@@ -32,17 +32,19 @@
*/
#define DEBUG_SUBSYSTEM S_RPC
+#include
#include
#ifdef HAVE_UIDGID_HEADER
#include
#endif
+#include
#include
#include
#include
char obd_jobid_var[JOBSTATS_JOBID_VAR_MAX_LEN + 1] = JOBSTATS_DISABLE;
-char obd_jobid_node[LUSTRE_JOBID_SIZE + 1];
+char obd_jobid_name[LUSTRE_JOBID_SIZE] = "%e.%u";
/* Get jobid of current process from stored variable or calculate
* it from pid and user_id.
@@ -52,9 +54,89 @@
* This is now deprecated.
*/
-int lustre_get_jobid(char *jobid)
+/*
+ * jobid_interpret_string()
+ *
+ * Interpret the jobfmt string to expand specified fields, like coredumps do:
+ * %e = executable
+ * %g = gid
+ * %h = hostname
+ * %j = jobid from environment
+ * %p = pid
+ * %u = uid
+ *
+ * Unknown escape strings are dropped. Other characters are copied through,
+ * excluding whitespace (to avoid making jobid parsing difficult).
+ *
+ * Return: -EOVERFLOW if the expanded string does not fit within @joblen
+ * 0 for success
+ */
+static int jobid_interpret_string(const char *jobfmt, char *jobid,
+ ssize_t joblen)
+{
+ char c;
+
+ while ((c = *jobfmt++) && joblen > 1) {
+ char f;
+ int l;
+
+ if (isspace(c)) /* Don't allow embedded spaces */
+ continue;
+
+ if (c != '%') {
+ *jobid = c;
+ joblen--;
+ jobid++;
+ continue;
+ }
+
+ switch ((f = *jobfmt++)) {
+ case 'e': /* executable name */
+ l = snprintf(jobid, joblen, "%s", current->comm);
+ break;
+ case 'g': /* group ID */
+ l = snprintf(jobid, joblen, "%u",
+ from_kgid(&init_user_ns, current_fsgid()));
+ break;
+ case 'h': /* hostname */
+ l = snprintf(jobid, joblen, "%s",
+ init_utsname()->nodename);
+ break;
+ case 'j': /* jobid requested by process
+ * - currently not supported
+ */
+ l = snprintf(jobid, joblen, "%s", "jobid");
+ break;
+ case 'p': /* process ID */
+ l = snprintf(jobid, joblen, "%u", current->pid);
+ break;
+ case 'u': /* user ID */
+ l = snprintf(jobid, joblen, "%u",
+ from_kuid(&init_user_ns, current_fsuid()));
+ break;
+ case '\0': /* '%' at end of format string */
+ l = 0;
+ goto out;
+ default: /* drop unknown %x format strings */
+ l = 0;
+ break;
+ }
+ jobid += l;
+ joblen -= l;
+ }
+ /*
+ * This points at the end of the buffer, so long as jobid is always
+ * incremented the same amount as joblen is decremented.
+ */
+out:
+ jobid[joblen - 1] = '\0';
+
+ return joblen < 0 ? -EOVERFLOW : 0;
+}
+
+int lustre_get_jobid(char *jobid, size_t joblen)
{
- char tmp_jobid[LUSTRE_JOBID_SIZE] = { 0 };
+ char tmp_jobid[LUSTRE_JOBID_SIZE] = "";
/* Jobstats isn't enabled */
if (strcmp(obd_jobid_var, JOBSTATS_DISABLE) == 0)
@@ -70,10 +152,11 @@ int lustre_get_jobid(char *jobid)
/* Whole node dedicated to single job */
if (strcmp(obd_jobid_var, JOBSTATS_NODELOCAL) == 0) {
- strcpy(tmp_jobid, obd_jobid_node);
- goto out_cache_jobid;
+ int rc2 = jobid_interpret_string(obd_jobid_name,
+ tmp_jobid, joblen);
+ if (!rc2)
+ goto out_cache_jobid;
}
-
return -ENOENT;
out_cache_jobid:
diff --git a/fs/lustre/obdclass/obd_sysfs.c b/fs/lustre/obdclass/obd_sysfs.c
index bac8e7c5..cd2917e 100644
--- a/fs/lustre/obdclass/obd_sysfs.c
+++ b/fs/lustre/obdclass/obd_sysfs.c
@@ -233,7 +233,7 @@ static ssize_t jobid_var_store(struct kobject *kobj, struct attribute *attr,
static ssize_t jobid_name_show(struct kobject *kobj, struct attribute *attr,
char *buf)
{
- return snprintf(buf, PAGE_SIZE, "%s\n", obd_jobid_node);
+ return snprintf(buf, PAGE_SIZE, "%s\n", obd_jobid_name);
}
static ssize_t jobid_name_store(struct kobject *kobj, struct attribute *attr,
@@ -243,13 +243,13 @@ static ssize_t jobid_name_store(struct kobject *kobj, struct attribute *attr,
if (!count || count > LUSTRE_JOBID_SIZE)
return -EINVAL;
- memcpy(obd_jobid_node, buffer, count);
+ memcpy(obd_jobid_name, buffer, count);
- obd_jobid_node[count] = 0;
+ obd_jobid_name[count] = 0;
/* Trim the trailing '\n' if any */
- if (obd_jobid_node[count - 1] == '\n')
- obd_jobid_node[count - 1] = 0;
+ if (obd_jobid_name[count - 1] == '\n')
+ obd_jobid_name[count - 1] = 0;
return count;
}
diff --git a/fs/lustre/ptlrpc/pack_generic.c b/fs/lustre/ptlrpc/pack_generic.c
index b6a4fd8..bc5e513 100644
--- a/fs/lustre/ptlrpc/pack_generic.c
+++ b/fs/lustre/ptlrpc/pack_generic.c
@@ -1406,9 +1406,9 @@ void lustre_msg_set_jobid(struct lustre_msg *msg, char *jobid)
LASSERTF(pb, "invalid msg %p: no ptlrpc body!\n", msg);
if (jobid)
- memcpy(pb->pb_jobid, jobid, LUSTRE_JOBID_SIZE);
+ memcpy(pb->pb_jobid, jobid, sizeof(pb->pb_jobid));
else if (pb->pb_jobid[0] == '\0')
- lustre_get_jobid(pb->pb_jobid);
+ lustre_get_jobid(pb->pb_jobid, sizeof(pb->pb_jobid));
return;
}
default:
diff --git a/include/uapi/linux/lustre/lustre_idl.h b/include/uapi/linux/lustre/lustre_idl.h
index 401f7ef..4e1605a2 100644
--- a/include/uapi/linux/lustre/lustre_idl.h
+++ b/include/uapi/linux/lustre/lustre_idl.h
@@ -635,7 +635,7 @@ struct ptlrpc_body_v3 {
__u64 pb_padding64_0;
__u64 pb_padding64_1;
__u64 pb_padding64_2;
- char pb_jobid[LUSTRE_JOBID_SIZE]; /* req: ASCII MPI jobid from env */
+ char pb_jobid[LUSTRE_JOBID_SIZE]; /* req: ASCII jobid from env + NUL */
};
#define ptlrpc_body ptlrpc_body_v3
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:12 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:12 -0500
Subject: [lustre-devel] [PATCH 024/622] lustre: llite: rename FSFILT_IOC_*
to system flags
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-25-git-send-email-jsimmons@infradead.org>
From: Jinshan Xiong
Those definitions were probably created for compatibility. Now that
FS_IOC_* have been existing in kernel for long time, we should use
them to avoid confusion.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10779
Lustre-commit: 7e3fc106d6e7 ("LU-10779 llite: rename FSFILT_IOC_* to system flags")
Signed-off-by: Jinshan Xiong
Reviewed-on: https://review.whamcloud.com/31546
Reviewed-by: James Simmons
Reviewed-by: Andreas Dilger
Signed-off-by: James Simmons
---
fs/lustre/llite/dir.c | 13 +++++++------
fs/lustre/llite/file.c | 19 ++++++++++---------
fs/lustre/llite/llite_lib.c | 4 ++--
3 files changed, 19 insertions(+), 17 deletions(-)
diff --git a/fs/lustre/llite/dir.c b/fs/lustre/llite/dir.c
index f21727b..b006e32 100644
--- a/fs/lustre/llite/dir.c
+++ b/fs/lustre/llite/dir.c
@@ -1108,18 +1108,19 @@ static long ll_dir_ioctl(struct file *file, unsigned int cmd, unsigned long arg)
ll_stats_ops_tally(ll_i2sbi(inode), LPROC_LL_IOCTL, 1);
switch (cmd) {
- case FSFILT_IOC_GETFLAGS:
- case FSFILT_IOC_SETFLAGS:
+ case FS_IOC_GETFLAGS:
+ case FS_IOC_SETFLAGS:
return ll_iocontrol(inode, file, cmd, arg);
- case FSFILT_IOC_GETVERSION_OLD:
case FSFILT_IOC_GETVERSION:
+ case FS_IOC_GETVERSION:
return put_user(inode->i_generation, (int __user *)arg);
/* We need to special case any other ioctls we want to handle,
* to send them to the MDS/OST as appropriate and to properly
* network encode the arg field.
- case FSFILT_IOC_SETVERSION_OLD:
- case FSFILT_IOC_SETVERSION:
- */
+ */
+ case FS_IOC_SETVERSION:
+ return -ENOTSUPP;
+
case LL_IOC_GET_MDTIDX: {
int mdtidx;
diff --git a/fs/lustre/llite/file.c b/fs/lustre/llite/file.c
index fe965b1..c3fb104b 100644
--- a/fs/lustre/llite/file.c
+++ b/fs/lustre/llite/file.c
@@ -3055,12 +3055,19 @@ static long ll_file_set_lease(struct file *file, struct ll_ioc_lease *ioc,
case LL_IOC_LOV_GETSTRIPE:
case LL_IOC_LOV_GETSTRIPE_NEW:
return ll_file_getstripe(inode, (void __user *)arg, 0);
- case FSFILT_IOC_GETFLAGS:
- case FSFILT_IOC_SETFLAGS:
+ case FS_IOC_GETFLAGS:
+ case FS_IOC_SETFLAGS:
return ll_iocontrol(inode, file, cmd, arg);
- case FSFILT_IOC_GETVERSION_OLD:
case FSFILT_IOC_GETVERSION:
+ case FS_IOC_GETVERSION:
return put_user(inode->i_generation, (int __user *)arg);
+ /* We need to special case any other ioctls we want to handle,
+ * to send them to the MDS/OST as appropriate and to properly
+ * network encode the arg field.
+ */
+ case FS_IOC_SETVERSION:
+ return -ENOTSUPP;
+
case LL_IOC_GROUP_LOCK:
return ll_get_grouplock(inode, file, arg);
case LL_IOC_GROUP_UNLOCK:
@@ -3068,12 +3075,6 @@ static long ll_file_set_lease(struct file *file, struct ll_ioc_lease *ioc,
case IOC_OBD_STATFS:
return ll_obd_statfs(inode, (void __user *)arg);
- /* We need to special case any other ioctls we want to handle,
- * to send them to the MDS/OST as appropriate and to properly
- * network encode the arg field.
- case FSFILT_IOC_SETVERSION_OLD:
- case FSFILT_IOC_SETVERSION:
- */
case LL_IOC_FLUSHCTX:
return ll_flush_ctx(inode);
case LL_IOC_PATH2FID: {
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index 7580d57..e2c7a4d 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -2037,7 +2037,7 @@ int ll_iocontrol(struct inode *inode, struct file *file,
int rc, flags = 0;
switch (cmd) {
- case FSFILT_IOC_GETFLAGS: {
+ case FS_IOC_GETFLAGS: {
struct mdt_body *body;
struct md_op_data *op_data;
@@ -2065,7 +2065,7 @@ int ll_iocontrol(struct inode *inode, struct file *file,
return put_user(flags, (int __user *)arg);
}
- case FSFILT_IOC_SETFLAGS: {
+ case FS_IOC_SETFLAGS: {
struct md_op_data *op_data;
struct cl_object *obj;
struct iattr *attr;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:02 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:02 -0500
Subject: [lustre-devel] [PATCH 014/622] lustre: mdc: fix possible NULL
pointer dereference
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-15-git-send-email-jsimmons@infradead.org>
From: Andreas Dilger
Fix two static analysis errors.
fs/lustre/mdc/mdc_dev.c: in mdc_enqueue_send(), pointer 'matched' return
from call to function 'ldlm_handle2lock' at line 704 may be NULL
and will be dereferenced at line 705.
If client is evicted between ldlm_lock_match() and ldlm_handle2lock()
the lock pointer could be NULL.
fs/lustre/lov/lov_dev.c:488 in lov_process_config, sscanf format
specification '%d' expects type 'int' for 'd', but parameter 3
has a different type '__u32'.
Converting to kstrtou32() requires changing the "index" variable type
from __u32 to u32, which is fine since it is only used internally,
fix up the few functions that are also passing "__u32 index" and the
resulting checkpatch.pl warnings.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10264
Lustre-commit: b89206476174 ("LU-10264 mdc: fix possible NULL pointer dereference")
Signed-off-by: Andreas Dilger
Reviewed-on: https://review.whamcloud.com/31621
Reviewed-by: Dmitry Eremin
Reviewed-by: Bob Glossman
Reviewed-by: James Simmons
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/lov/lov_obd.c | 45 ++++++++++++++++++++++++---------------------
fs/lustre/mdc/mdc_dev.c | 2 +-
2 files changed, 25 insertions(+), 22 deletions(-)
diff --git a/fs/lustre/lov/lov_obd.c b/fs/lustre/lov/lov_obd.c
index 1708fa9..26637bc 100644
--- a/fs/lustre/lov/lov_obd.c
+++ b/fs/lustre/lov/lov_obd.c
@@ -312,7 +312,8 @@ static int lov_disconnect(struct obd_export *exp)
{
struct obd_device *obd = class_exp2obd(exp);
struct lov_obd *lov = &obd->u.lov;
- int i, rc;
+ u32 index;
+ int rc;
if (!lov->lov_tgts)
goto out;
@@ -321,19 +322,19 @@ static int lov_disconnect(struct obd_export *exp)
lov->lov_connects--;
if (lov->lov_connects != 0) {
/* why should there be more than 1 connect? */
- CERROR("disconnect #%d\n", lov->lov_connects);
+ CWARN("%s: unexpected disconnect #%d\n",
+ obd->obd_name, lov->lov_connects);
goto out;
}
- /* Let's hold another reference so lov_del_obd doesn't spin through
- * putref every time
- */
+ /* hold another ref so lov_del_obd() doesn't spin in putref each time */
lov_tgts_getref(obd);
- for (i = 0; i < lov->desc.ld_tgt_count; i++) {
- if (lov->lov_tgts[i] && lov->lov_tgts[i]->ltd_exp) {
- /* Disconnection is the last we know about an obd */
- lov_del_target(obd, i, NULL, lov->lov_tgts[i]->ltd_gen);
+ for (index = 0; index < lov->desc.ld_tgt_count; index++) {
+ if (lov->lov_tgts[index] && lov->lov_tgts[index]->ltd_exp) {
+ /* Disconnection is the last we know about an OBD */
+ lov_del_target(obd, index, NULL,
+ lov->lov_tgts[index]->ltd_gen);
}
}
@@ -490,13 +491,12 @@ static int lov_add_target(struct obd_device *obd, struct obd_uuid *uuidp,
uuidp->uuid, index, gen, active);
if (gen <= 0) {
- CERROR("request to add OBD %s with invalid generation: %d\n",
- uuidp->uuid, gen);
+ CERROR("%s: request to add '%s' with invalid generation: %d\n",
+ obd->obd_name, uuidp->uuid, gen);
return -EINVAL;
}
- tgt_obd = class_find_client_obd(uuidp, LUSTRE_OSC_NAME,
- &obd->obd_uuid);
+ tgt_obd = class_find_client_obd(uuidp, LUSTRE_OSC_NAME, &obd->obd_uuid);
if (!tgt_obd)
return -EINVAL;
@@ -504,10 +504,11 @@ static int lov_add_target(struct obd_device *obd, struct obd_uuid *uuidp,
if ((index < lov->lov_tgt_size) && lov->lov_tgts[index]) {
tgt = lov->lov_tgts[index];
- CERROR("UUID %s already assigned at LOV target index %d\n",
- obd_uuid2str(&tgt->ltd_uuid), index);
+ rc = -EEXIST;
+ CERROR("%s: UUID %s already assigned at index %d: rc = %d\n",
+ obd->obd_name, obd_uuid2str(&tgt->ltd_uuid), index, rc);
mutex_unlock(&lov->lov_lock);
- return -EEXIST;
+ return rc;
}
if (index >= lov->lov_tgt_size) {
@@ -602,8 +603,8 @@ static int lov_add_target(struct obd_device *obd, struct obd_uuid *uuidp,
out:
if (rc) {
- CERROR("add failed (%d), deleting %s\n", rc,
- obd_uuid2str(&tgt->ltd_uuid));
+ CERROR("%s: add failed, deleting %s: rc = %d\n",
+ obd->obd_name, obd_uuid2str(&tgt->ltd_uuid), rc);
lov_del_target(obd, index, NULL, 0);
}
lov_tgts_putref(obd);
@@ -860,6 +861,7 @@ int lov_process_config_base(struct obd_device *obd, struct lustre_cfg *lcfg,
case LCFG_LOV_DEL_OBD: {
u32 index;
int gen;
+
/* lov_modify_tgts add 0:lov_mdsA 1:ost1_UUID 2:0 3:1 */
if (LUSTRE_CFG_BUFLEN(lcfg, 1) > sizeof(obd_uuid.uuid)) {
rc = -EINVAL;
@@ -868,11 +870,11 @@ int lov_process_config_base(struct obd_device *obd, struct lustre_cfg *lcfg,
obd_str2uuid(&obd_uuid, lustre_cfg_buf(lcfg, 1));
- rc = kstrtoint(lustre_cfg_buf(lcfg, 2), 10, indexp);
- if (rc < 0)
+ rc = kstrtou32(lustre_cfg_buf(lcfg, 2), 10, indexp);
+ if (rc)
goto out;
rc = kstrtoint(lustre_cfg_buf(lcfg, 3), 10, genp);
- if (rc < 0)
+ if (rc)
goto out;
index = *indexp;
gen = *genp;
@@ -882,6 +884,7 @@ int lov_process_config_base(struct obd_device *obd, struct lustre_cfg *lcfg,
rc = lov_add_target(obd, &obd_uuid, index, gen, 0);
else
rc = lov_del_target(obd, index, &obd_uuid, gen);
+
goto out;
}
case LCFG_PARAM: {
diff --git a/fs/lustre/mdc/mdc_dev.c b/fs/lustre/mdc/mdc_dev.c
index ca0822d..80e3120 100644
--- a/fs/lustre/mdc/mdc_dev.c
+++ b/fs/lustre/mdc/mdc_dev.c
@@ -684,7 +684,7 @@ int mdc_enqueue_send(const struct lu_env *env, struct obd_export *exp,
return ELDLM_OK;
matched = ldlm_handle2lock(&lockh);
- if (ldlm_is_kms_ignore(matched))
+ if (!matched || ldlm_is_kms_ignore(matched))
goto no_match;
if (mdc_set_dom_lock_data(env, matched, einfo->ei_cbdata)) {
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:14 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:14 -0500
Subject: [lustre-devel] [PATCH 026/622] lustre: ptlrpc: fix
test_req_buffer_pressure behavior
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-27-git-send-email-jsimmons@infradead.org>
From: Bruno Faccini
In 2nd patch for LU-9372, to allow limiting number of rqbd-buffers,
a wrong and unnecessary test had been added to enhance
test_req_buffer_pressure feature.
This patch fixes this issue by removing such test.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10826
Lustre-commit: 040eca67f8d5 ("LU-10826 ptlrpc: fix test_req_buffer_pressure behavior")
Signed-off-by: Bruno Faccini
Reviewed-on: https://review.whamcloud.com/31690
Reviewed-by: Wang Shilong
Reviewed-by: Li Dongyang
Reviewed-by: Dmitry Eremin
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/ptlrpc/service.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/fs/lustre/ptlrpc/service.c b/fs/lustre/ptlrpc/service.c
index 3c61e83..8dae21a 100644
--- a/fs/lustre/ptlrpc/service.c
+++ b/fs/lustre/ptlrpc/service.c
@@ -150,8 +150,7 @@
/* NB: another thread might have recycled enough rqbds, we
* need to make sure it wouldn't over-allocate, see LU-1212.
*/
- if (test_req_buffer_pressure ||
- svcpt->scp_nrqbds_posted >= svc->srv_nbuf_per_group ||
+ if (svcpt->scp_nrqbds_posted >= svc->srv_nbuf_per_group ||
(svc->srv_nrqbds_max != 0 &&
svcpt->scp_nrqbds_total > svc->srv_nrqbds_max))
break;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:17 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:17 -0500
Subject: [lustre-devel] [PATCH 029/622] lustre: mdc: resend quotactl if
needed
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-30-git-send-email-jsimmons@infradead.org>
From: Hongchao Zhang
In mdc_quotactl, it is better to resend the quotactl request
if reconnection or failover is triggered during the process.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10368
Lustre-commit: d511918e8eb7 ("LU-10368 mdc: resend quotactl if needed")
Signed-off-by: Hongchao Zhang
Reviewed-on: https://review.whamcloud.com/31773
Reviewed-by: Fan Yong
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/mdc/mdc_request.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/fs/lustre/mdc/mdc_request.c b/fs/lustre/mdc/mdc_request.c
index 5718db2..feac374 100644
--- a/fs/lustre/mdc/mdc_request.c
+++ b/fs/lustre/mdc/mdc_request.c
@@ -1867,7 +1867,7 @@ static int mdc_ioc_hsm_ct_start(struct obd_export *exp,
struct lustre_kernelcomm *lk);
static int mdc_quotactl(struct obd_device *unused, struct obd_export *exp,
- struct obd_quotactl *oqctl)
+ struct obd_quotactl *oqctl)
{
struct ptlrpc_request *req;
struct obd_quotactl *oqc;
@@ -1884,7 +1884,6 @@ static int mdc_quotactl(struct obd_device *unused, struct obd_export *exp,
ptlrpc_request_set_replen(req);
ptlrpc_at_set_req_timeout(req);
- req->rq_no_resend = 1;
rc = ptlrpc_queue_wait(req);
if (rc)
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:29 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:29 -0500
Subject: [lustre-devel] [PATCH 041/622] lustre: lmv: dir page is released
while in use
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-42-git-send-email-jsimmons@infradead.org>
From: Lai Siyao
When popping stripe dirent, if it reaches page end,
stripe_dirent_next() releases current page and then reads next one,
but current dirent is still in use, as will cause wrong values used,
and trigger assertion.
This patch changes to not read next page upon reaching end, but
leave it to next dirent read.
WC-bug-id: https://jira.whamcloud.com/browse/LU-9857
Lustre-commit: b51e8d6b53a3 ("LU-9857 lmv: dir page is released while in use")
Signed-off-by: Lai Siyao
Reviewed-on: https://review.whamcloud.com/32180
Reviewed-by: Fan Yong
Reviewed-by: John L. Hammond
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/lmv/lmv_obd.c | 123 +++++++++++++++++++++++-------------------------
1 file changed, 60 insertions(+), 63 deletions(-)
diff --git a/fs/lustre/lmv/lmv_obd.c b/fs/lustre/lmv/lmv_obd.c
index d0f626f..c7bf8c7 100644
--- a/fs/lustre/lmv/lmv_obd.c
+++ b/fs/lustre/lmv/lmv_obd.c
@@ -2016,7 +2016,7 @@ struct lmv_dir_ctxt {
struct stripe_dirent ldc_stripes[0];
};
-static inline void put_stripe_dirent(struct stripe_dirent *stripe)
+static inline void stripe_dirent_unload(struct stripe_dirent *stripe)
{
if (stripe->sd_page) {
kunmap(stripe->sd_page);
@@ -2031,62 +2031,77 @@ static inline void put_lmv_dir_ctxt(struct lmv_dir_ctxt *ctxt)
int i;
for (i = 0; i < ctxt->ldc_count; i++)
- put_stripe_dirent(&ctxt->ldc_stripes[i]);
+ stripe_dirent_unload(&ctxt->ldc_stripes[i]);
}
-static struct lu_dirent *stripe_dirent_next(struct lmv_dir_ctxt *ctxt,
+/* if @ent is dummy, or . .., get next */
+static struct lu_dirent *stripe_dirent_get(struct lmv_dir_ctxt *ctxt,
+ struct lu_dirent *ent,
+ int stripe_index)
+{
+ for (; ent; ent = lu_dirent_next(ent)) {
+ /* Skip dummy entry */
+ if (le16_to_cpu(ent->lde_namelen) == 0)
+ continue;
+
+ /* skip . and .. for other stripes */
+ if (stripe_index &&
+ (strncmp(ent->lde_name, ".",
+ le16_to_cpu(ent->lde_namelen)) == 0 ||
+ strncmp(ent->lde_name, "..",
+ le16_to_cpu(ent->lde_namelen)) == 0))
+ continue;
+
+ if (le64_to_cpu(ent->lde_hash) >= ctxt->ldc_hash)
+ break;
+ }
+
+ return ent;
+}
+
+static struct lu_dirent *stripe_dirent_load(struct lmv_dir_ctxt *ctxt,
struct stripe_dirent *stripe,
int stripe_index)
{
+ struct md_op_data *op_data = ctxt->ldc_op_data;
+ struct lmv_oinfo *oinfo;
+ struct lu_fid fid = op_data->op_fid1;
+ struct inode *inode = op_data->op_data;
+ struct lmv_tgt_desc *tgt;
struct lu_dirent *ent = stripe->sd_ent;
u64 hash = ctxt->ldc_hash;
- u64 end;
int rc = 0;
LASSERT(stripe == &ctxt->ldc_stripes[stripe_index]);
-
- if (stripe->sd_eof)
- return NULL;
-
- if (ent) {
- ent = lu_dirent_next(ent);
- if (!ent) {
-check_eof:
- end = le64_to_cpu(stripe->sd_dp->ldp_hash_end);
-
- LASSERTF(hash <= end, "hash %llx end %llx\n",
- hash, end);
+ LASSERT(!ent);
+
+ do {
+ if (stripe->sd_page) {
+ u64 end = le64_to_cpu(stripe->sd_dp->ldp_hash_end);
+
+ /* @hash should be the last dirent hash */
+ LASSERTF(hash <= end,
+ "ctxt@%p stripe@%p hash %llx end %llx\n",
+ ctxt, stripe, hash, end);
+ /* unload last page */
+ stripe_dirent_unload(stripe);
+ /* eof */
if (end == MDS_DIR_END_OFF) {
stripe->sd_ent = NULL;
stripe->sd_eof = true;
- return NULL;
+ break;
}
-
- put_stripe_dirent(stripe);
hash = end;
}
- }
-
- if (!ent) {
- struct md_op_data *op_data = ctxt->ldc_op_data;
- struct lmv_oinfo *oinfo;
- struct lu_fid fid = op_data->op_fid1;
- struct inode *inode = op_data->op_data;
- struct lmv_tgt_desc *tgt;
-
- LASSERT(!stripe->sd_page);
oinfo = &op_data->op_mea1->lsm_md_oinfo[stripe_index];
tgt = lmv_get_target(ctxt->ldc_lmv, oinfo->lmo_mds, NULL);
if (IS_ERR(tgt)) {
rc = PTR_ERR(tgt);
- goto out;
+ break;
}
- /*
- * op_data will be shared by each stripe, so we need
- * reset these value for each stripe
- */
+ /* op_data is shared by stripes, reset after use */
op_data->op_fid1 = oinfo->lmo_fid;
op_data->op_fid2 = oinfo->lmo_fid;
op_data->op_data = oinfo->lmo_root;
@@ -2099,42 +2114,24 @@ static struct lu_dirent *stripe_dirent_next(struct lmv_dir_ctxt *ctxt,
op_data->op_data = inode;
if (rc)
- goto out;
-
- stripe->sd_dp = page_address(stripe->sd_page);
- ent = lu_dirent_start(stripe->sd_dp);
- }
-
- for (; ent; ent = lu_dirent_next(ent)) {
- /* Skip dummy entry */
- if (!le16_to_cpu(ent->lde_namelen))
- continue;
-
- /* skip . and .. for other stripes */
- if (stripe_index &&
- (strncmp(ent->lde_name, ".",
- le16_to_cpu(ent->lde_namelen)) == 0 ||
- strncmp(ent->lde_name, "..",
- le16_to_cpu(ent->lde_namelen)) == 0))
- continue;
-
- if (le64_to_cpu(ent->lde_hash) >= hash)
break;
- }
- if (!ent)
- goto check_eof;
+ stripe->sd_dp = page_address(stripe->sd_page);
+ ent = stripe_dirent_get(ctxt, lu_dirent_start(stripe->sd_dp),
+ stripe_index);
+ /* in case a page filled with ., .. and dummy, read next */
+ } while (!ent);
-out:
stripe->sd_ent = ent;
- /* treat error as eof, so dir can be partially accessed */
if (rc) {
- put_stripe_dirent(stripe);
+ LASSERT(!ent);
+ /* treat error as eof, so dir can be partially accessed */
stripe->sd_eof = true;
LCONSOLE_WARN("dir " DFID " stripe %d readdir failed: %d, directory is partially accessed!\n",
PFID(&ctxt->ldc_op_data->op_fid1), stripe_index,
rc);
}
+
return ent;
}
@@ -2186,8 +2183,7 @@ static struct lu_dirent *lmv_dirent_next(struct lmv_dir_ctxt *ctxt)
continue;
if (!stripe->sd_ent) {
- /* locate starting entry */
- stripe_dirent_next(ctxt, stripe, i);
+ stripe_dirent_load(ctxt, stripe, i);
if (!stripe->sd_ent) {
LASSERT(stripe->sd_eof);
continue;
@@ -2208,7 +2204,8 @@ static struct lu_dirent *lmv_dirent_next(struct lmv_dir_ctxt *ctxt)
stripe = &ctxt->ldc_stripes[min];
ent = stripe->sd_ent;
/* pop found dirent */
- stripe_dirent_next(ctxt, stripe, min);
+ stripe->sd_ent = stripe_dirent_get(ctxt, lu_dirent_next(ent),
+ min);
}
return ent;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:43 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:43 -0500
Subject: [lustre-devel] [PATCH 055/622] lustre: ldlm: handle lock converts
in cancel handler
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-56-git-send-email-jsimmons@infradead.org>
From: Mikhail Pershin
- Use cancel portals and high-priority handling for lock
converts. Update ldlm_cancel_handler to understand
LDLM_CONVERT RPC for that.
- Use ns_dirty_age_limit for lock convert - don't convert too old
locks.
- Check for empty converts and skip such
WC-bug-id: https://jira.whamcloud.com/browse/LU-10175
Lustre-commit: 541902a3f934 ("LU-10175 ldlm: handle lock converts in cancel handler")
Signed-off-by: Mikhail Pershin
Reviewed-on: https://review.whamcloud.com/32314
Reviewed-by: Fan Yong
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/include/lustre_export.h | 6 ++++++
fs/lustre/ldlm/ldlm_inodebits.c | 19 ++++++++++++++-----
fs/lustre/ldlm/ldlm_request.c | 39 +++++++++++++++++++++++++++++++--------
fs/lustre/llite/llite_lib.c | 2 +-
fs/lustre/llite/namei.c | 7 ++++++-
5 files changed, 58 insertions(+), 15 deletions(-)
diff --git a/fs/lustre/include/lustre_export.h b/fs/lustre/include/lustre_export.h
index de3b109..57cf68b 100644
--- a/fs/lustre/include/lustre_export.h
+++ b/fs/lustre/include/lustre_export.h
@@ -269,9 +269,15 @@ static inline int exp_connect_flr(struct obd_export *exp)
return !!(exp_connect_flags2(exp) & OBD_CONNECT2_FLR);
}
+static inline int exp_connect_lock_convert(struct obd_export *exp)
+{
+ return !!(exp_connect_flags2(exp) & OBD_CONNECT2_LOCK_CONVERT);
+}
+
struct obd_export *class_conn2export(struct lustre_handle *conn);
#define KKUC_CT_DATA_MAGIC 0x092013cea
+
struct kkuc_ct_data {
u32 kcd_magic;
u32 kcd_archive;
diff --git a/fs/lustre/ldlm/ldlm_inodebits.c b/fs/lustre/ldlm/ldlm_inodebits.c
index ddbf8d4..9cf3c5f 100644
--- a/fs/lustre/ldlm/ldlm_inodebits.c
+++ b/fs/lustre/ldlm/ldlm_inodebits.c
@@ -81,7 +81,7 @@ int ldlm_inodebits_drop(struct ldlm_lock *lock, u64 to_drop)
/* Just return if there are no conflicting bits */
if ((lock->l_policy_data.l_inodebits.bits & to_drop) == 0) {
- LDLM_WARN(lock, "try to drop unset bits %#llx/%#llx\n",
+ LDLM_WARN(lock, "try to drop unset bits %#llx/%#llx",
lock->l_policy_data.l_inodebits.bits, to_drop);
/* nothing to do */
return 0;
@@ -111,7 +111,7 @@ int ldlm_cli_dropbits(struct ldlm_lock *lock, u64 drop_bits)
ldlm_lock2handle(lock, &lockh);
lock_res_and_lock(lock);
- /* check if all bits are cancelled */
+ /* check if all bits are blocked */
if (!(lock->l_policy_data.l_inodebits.bits & ~drop_bits)) {
unlock_res_and_lock(lock);
/* return error to continue with cancel */
@@ -119,6 +119,13 @@ int ldlm_cli_dropbits(struct ldlm_lock *lock, u64 drop_bits)
goto exit;
}
+ /* check if no common bits, consider this as successful convert */
+ if (!(lock->l_policy_data.l_inodebits.bits & drop_bits)) {
+ unlock_res_and_lock(lock);
+ rc = 0;
+ goto exit;
+ }
+
/* check if there is race with cancel */
if (ldlm_is_canceling(lock) || ldlm_is_cancel(lock)) {
unlock_res_and_lock(lock);
@@ -167,9 +174,11 @@ int ldlm_cli_dropbits(struct ldlm_lock *lock, u64 drop_bits)
rc = ldlm_cli_convert(lock, &flags);
if (rc) {
lock_res_and_lock(lock);
- ldlm_clear_converting(lock);
- ldlm_set_cbpending(lock);
- ldlm_set_bl_ast(lock);
+ if (ldlm_is_converting(lock)) {
+ ldlm_clear_converting(lock);
+ ldlm_set_cbpending(lock);
+ ldlm_set_bl_ast(lock);
+ }
unlock_res_and_lock(lock);
goto exit;
}
diff --git a/fs/lustre/ldlm/ldlm_request.c b/fs/lustre/ldlm/ldlm_request.c
index 5833f59..ad54bd2 100644
--- a/fs/lustre/ldlm/ldlm_request.c
+++ b/fs/lustre/ldlm/ldlm_request.c
@@ -854,7 +854,7 @@ static int lock_convert_interpret(const struct lu_env *env,
aa->lock_handle.cookie, reply->lock_handle.cookie,
req->rq_export->exp_client_uuid.uuid,
libcfs_id2str(req->rq_peer));
- rc = -ESTALE;
+ rc = ELDLM_NO_LOCK_DATA;
goto out;
}
@@ -905,15 +905,30 @@ static int lock_convert_interpret(const struct lu_env *env,
unlock_res_and_lock(lock);
out:
if (rc) {
+ int flag;
+
lock_res_and_lock(lock);
if (ldlm_is_converting(lock)) {
ldlm_clear_converting(lock);
ldlm_set_cbpending(lock);
ldlm_set_bl_ast(lock);
+ lock->l_policy_data.l_inodebits.cancel_bits = 0;
}
unlock_res_and_lock(lock);
- }
+ /* fallback to normal lock cancel. If rc means there is no
+ * valid lock on server, do only local cancel
+ */
+ if (rc == ELDLM_NO_LOCK_DATA)
+ flag = LCF_LOCAL;
+ else
+ flag = LCF_ASYNC;
+
+ rc = ldlm_cli_cancel(&aa->lock_handle, flag);
+ if (rc < 0)
+ LDLM_DEBUG(lock, "failed to cancel lock: rc = %d\n",
+ rc);
+ }
LDLM_LOCK_PUT(lock);
return rc;
}
@@ -942,6 +957,15 @@ int ldlm_cli_convert(struct ldlm_lock *lock, u32 *flags)
return -EINVAL;
}
+ /* this is better to check earlier and it is done so already,
+ * but this check is kept too as final one to issue an error
+ * if any new code will miss such check.
+ */
+ if (!exp_connect_lock_convert(exp)) {
+ LDLM_ERROR(lock, "server doesn't support lock convert\n");
+ return -EPROTO;
+ }
+
if (lock->l_resource->lr_type != LDLM_IBITS) {
LDLM_ERROR(lock, "convert works with IBITS locks only.");
return -EINVAL;
@@ -970,13 +994,12 @@ int ldlm_cli_convert(struct ldlm_lock *lock, u32 *flags)
ptlrpc_request_set_replen(req);
- /* That could be useful to use cancel portals for convert as well
- * as high-priority handling. This will require changes in
- * ldlm_cancel_handler to understand convert RPC as well.
- *
- * req->rq_request_portal = LDLM_CANCEL_REQUEST_PORTAL;
- * req->rq_reply_portal = LDLM_CANCEL_REPLY_PORTAL;
+ /*
+ * Use cancel portals for convert as well as high-priority handling.
*/
+ req->rq_request_portal = LDLM_CANCEL_REQUEST_PORTAL;
+ req->rq_reply_portal = LDLM_CANCEL_REPLY_PORTAL;
+
ptlrpc_at_set_req_timeout(req);
if (exp->exp_obd->obd_svc_stats)
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index dff349f..0844318 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -209,7 +209,7 @@ static int client_common_fill_super(struct super_block *sb, char *md, char *dt)
OBD_CONNECT_GRANT_PARAM |
OBD_CONNECT_SHORTIO | OBD_CONNECT_FLAGS2;
- data->ocd_connect_flags2 = OBD_CONNECT2_FLR;
+ data->ocd_connect_flags2 = OBD_CONNECT2_FLR | OBD_CONNECT2_LOCK_CONVERT;
if (sbi->ll_flags & LL_SBI_LRU_RESIZE)
data->ocd_connect_flags |= OBD_CONNECT_LRU_RESIZE;
diff --git a/fs/lustre/llite/namei.c b/fs/lustre/llite/namei.c
index 8b1a1ca..f835abb 100644
--- a/fs/lustre/llite/namei.c
+++ b/fs/lustre/llite/namei.c
@@ -371,11 +371,16 @@ void ll_lock_cancel_bits(struct ldlm_lock *lock, u64 to_cancel)
*/
int ll_md_need_convert(struct ldlm_lock *lock)
{
+ struct ldlm_namespace *ns = ldlm_lock_to_ns(lock);
struct inode *inode;
u64 wanted = lock->l_policy_data.l_inodebits.cancel_bits;
u64 bits = lock->l_policy_data.l_inodebits.bits & ~wanted;
enum ldlm_mode mode = LCK_MINMODE;
+ if (!lock->l_conn_export ||
+ !exp_connect_lock_convert(lock->l_conn_export))
+ return 0;
+
if (!wanted || !bits || ldlm_is_cancel(lock))
return 0;
@@ -410,7 +415,7 @@ int ll_md_need_convert(struct ldlm_lock *lock)
lock_res_and_lock(lock);
if (ktime_after(ktime_get(),
ktime_add(lock->l_last_used,
- ktime_set(10, 0)))) {
+ ktime_set(ns->ns_dirty_age_limit, 0)))) {
unlock_res_and_lock(lock);
return 0;
}
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:46 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:46 -0500
Subject: [lustre-devel] [PATCH 058/622] lustre: quota: add default quota
setting support
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-59-git-send-email-jsimmons@infradead.org>
From: Hongchao Zhang
Similar function which is motivated by GPFS which is friendly
feature for cluster administrators to manage quota.
Lazy Quota default setting support, here is basic idea:
Default quota setting is global quota setting for user, group,
project quotas, if default quota is set for one quota type,
newer created users/groups/projects will inherit this setting
automatically, since Lustre itself don't have ideas when new
users created, they could only know when this users trying to
acquire space from Lustre.
So we try to implement lazy quota setting inherit, Slave firstly
check if there exists default quota setting, if exists, it will
force slave to acquire quota from master, and master will detect
whether default quota is set, then it will set this quota and also
return proper grant space to slave.
To implement this and reuse existed quota APIs, we try to manage
the default quota in the quota record of 0 id, and enforce the
quota check when reading the quota recored from disk.
In the current Lustre implementation, the grace time is either
the time or the timestamp to be used after some quota ID exceeds
the soft limt, then 48bits should be enough for it, its high 16bits
can be used as kinds of quota flags, this patch will use one of
them as the default quota flag.
The global quota record used by default quota will set its soft
and hard limit as zero, its grace time will contain the default flag.
Use lfs setquota -U/-G/-P to set default quota.
Use lfs setquota -u/-g/-p foo -d to set foo to use default quota
Use lfs quota -U/-G/-P to show default quota.
WC-bug-id: https://jira.whamcloud.com/browse/LU-7816
Lustre-commit: 530881fe4ee2 ("LU-7816 quota: add default quota setting support")
Signed-off-by: Wang Shilong
Signed-off-by: Hongchao Zhang
Reviewed-on: https://review.whamcloud.com/32306
Reviewed-by: Fan Yong
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/dir.c | 4 +++-
include/uapi/linux/lustre/lustre_user.h | 22 ++++++++++++++++++++++
2 files changed, 25 insertions(+), 1 deletion(-)
diff --git a/fs/lustre/llite/dir.c b/fs/lustre/llite/dir.c
index b006e32..c0c3bf0 100644
--- a/fs/lustre/llite/dir.c
+++ b/fs/lustre/llite/dir.c
@@ -949,10 +949,12 @@ static int quotactl_ioctl(struct ll_sb_info *sbi, struct if_quotactl *qctl)
switch (cmd) {
case Q_SETQUOTA:
case Q_SETINFO:
+ case LUSTRE_Q_SETDEFAULT:
if (!capable(CAP_SYS_ADMIN))
return -EPERM;
break;
case Q_GETQUOTA:
+ case LUSTRE_Q_GETDEFAULT:
if (check_owner(type, id) && !capable(CAP_SYS_ADMIN))
return -EPERM;
break;
@@ -960,7 +962,7 @@ static int quotactl_ioctl(struct ll_sb_info *sbi, struct if_quotactl *qctl)
break;
default:
CERROR("unsupported quotactl op: %#x\n", cmd);
- return -ENOTTY;
+ return -ENOTSUPP;
}
if (valid != QC_GENERAL) {
diff --git a/include/uapi/linux/lustre/lustre_user.h b/include/uapi/linux/lustre/lustre_user.h
index 5405e1b..5956f33 100644
--- a/include/uapi/linux/lustre/lustre_user.h
+++ b/include/uapi/linux/lustre/lustre_user.h
@@ -728,6 +728,28 @@ static inline void obd_uuid2fsname(char *buf, char *uuid, int buflen)
/* lustre-specific control commands */
#define LUSTRE_Q_INVALIDATE 0x80000b /* deprecated as of 2.4 */
#define LUSTRE_Q_FINVALIDATE 0x80000c /* deprecated as of 2.4 */
+#define LUSTRE_Q_GETDEFAULT 0x80000d /* get default quota */
+#define LUSTRE_Q_SETDEFAULT 0x80000e /* set default quota */
+
+/* In the current Lustre implementation, the grace time is either the time
+ * or the timestamp to be used after some quota ID exceeds the soft limt,
+ * 48 bits should be enough, its high 16 bits can be used as quota flags.
+ */
+#define LQUOTA_GRACE_BITS 48
+#define LQUOTA_GRACE_MASK ((1ULL << LQUOTA_GRACE_BITS) - 1)
+#define LQUOTA_GRACE_MAX LQUOTA_GRACE_MASK
+#define LQUOTA_GRACE(t) (t & LQUOTA_GRACE_MASK)
+#define LQUOTA_FLAG(t) (t >> LQUOTA_GRACE_BITS)
+#define LQUOTA_GRACE_FLAG(t, f) ((__u64)t | (__u64)f << LQUOTA_GRACE_BITS)
+
+/* different quota flags */
+
+/* the default quota flag, the corresponding quota ID will use the default
+ * quota setting, the hardlimit and softlimit of its quota record in the global
+ * quota file will be set to 0, the low 48 bits of the grace will be set to 0
+ * and high 16 bits will contain this flag (see above comment).
+ */
+#define LQUOTA_FLAG_DEFAULT 0x0001
#define ALLQUOTA 255 /* set all quota */
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:09:06 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:09:06 -0500
Subject: [lustre-devel] [PATCH 078/622] lnet: add monitor thread
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-79-git-send-email-jsimmons@infradead.org>
From: Amir Shehata
Refactored the router checker thread to be the monitor thread.
The monitor thread will check router aliveness, expires messages
on the active list, recover local and remote NIs and resend messages.
In this patch it only checks router aliveness.
A deadline on the message is also added to keep track of when this
message should expire.
WC-bug-id: https://jira.whamcloud.com/browse/LU-9120
Lustre-commit: b01e6fce1c98 ("LU-9120 lnet: add monitor thread")
Signed-off-by: Amir Shehata
Reviewed-on: https://review.whamcloud.com/32763
Reviewed-by: Sonia Sharma
Reviewed-by: Olaf Weber
Reviewed-by: Chris Horn
Signed-off-by: James Simmons
---
include/linux/lnet/lib-lnet.h | 11 ++-
include/linux/lnet/lib-types.h | 27 +++----
net/lnet/lnet/api-ni.c | 12 ++--
net/lnet/lnet/lib-move.c | 98 ++++++++++++++++++++++++++
net/lnet/lnet/lib-msg.c | 9 ++-
net/lnet/lnet/router.c | 156 +++++++++++++----------------------------
6 files changed, 185 insertions(+), 128 deletions(-)
diff --git a/include/linux/lnet/lib-lnet.h b/include/linux/lnet/lib-lnet.h
index 5e13d32..2c3f665 100644
--- a/include/linux/lnet/lib-lnet.h
+++ b/include/linux/lnet/lib-lnet.h
@@ -714,8 +714,15 @@ int lnet_sock_connect(struct socket **sockp, int *fatal,
int lnet_peers_start_down(void);
int lnet_peer_buffer_credits(struct lnet_net *net);
-int lnet_router_checker_start(void);
-void lnet_router_checker_stop(void);
+int lnet_monitor_thr_start(void);
+void lnet_monitor_thr_stop(void);
+
+bool lnet_router_checker_active(void);
+void lnet_check_routers(void);
+int lnet_router_pre_mt_start(void);
+void lnet_router_post_mt_start(void);
+void lnet_prune_rc_data(int wait_unlink);
+void lnet_router_cleanup(void);
void lnet_router_ni_update_locked(struct lnet_peer_ni *gw, u32 net);
void lnet_swap_pinginfo(struct lnet_ping_buffer *pbuf);
diff --git a/include/linux/lnet/lib-types.h b/include/linux/lnet/lib-types.h
index 0ed325a..e1a56a1 100644
--- a/include/linux/lnet/lib-types.h
+++ b/include/linux/lnet/lib-types.h
@@ -79,6 +79,12 @@ struct lnet_msg {
lnet_nid_t msg_src_nid_param;
lnet_nid_t msg_rtr_nid_param;
+ /*
+ * Deadline for the message after which it will be finalized if it
+ * has not completed.
+ */
+ ktime_t msg_deadline;
+
/* committed for sending */
unsigned int msg_tx_committed:1;
/* CPT # this message committed for sending */
@@ -905,9 +911,9 @@ struct lnet_msg_container {
/* Router Checker states */
enum lnet_rc_state {
- LNET_RC_STATE_SHUTDOWN, /* not started */
- LNET_RC_STATE_RUNNING, /* started up OK */
- LNET_RC_STATE_STOPPING, /* telling thread to stop */
+ LNET_MT_STATE_SHUTDOWN, /* not started */
+ LNET_MT_STATE_RUNNING, /* started up OK */
+ LNET_MT_STATE_STOPPING, /* telling thread to stop */
};
/* LNet states */
@@ -1014,8 +1020,8 @@ struct lnet {
/* discovery startup/shutdown state */
int ln_dc_state;
- /* router checker startup/shutdown state */
- enum lnet_rc_state ln_rc_state;
+ /* monitor thread startup/shutdown state */
+ enum lnet_rc_state ln_mt_state;
/* router checker's event queue */
struct lnet_handle_eq ln_rc_eqh;
/* rcd still pending on net */
@@ -1023,7 +1029,7 @@ struct lnet {
/* rcd ready for free */
struct list_head ln_rcd_zombie;
/* serialise startup/shutdown */
- struct completion ln_rc_signal;
+ struct completion ln_mt_signal;
struct mutex ln_api_mutex;
struct mutex ln_lnd_mutex;
@@ -1053,13 +1059,10 @@ struct lnet {
*/
bool ln_nis_from_mod_params;
- /*
- * waitq for router checker. As long as there are no routes in
- * the list, the router checker will sleep on this queue. when
- * routes are added the thread will wake up
+ /* waitq for the monitor thread. The monitor thread takes care of
+ * checking routes, timedout messages and resending messages.
*/
- wait_queue_head_t ln_rc_waitq;
-
+ wait_queue_head_t ln_mt_waitq;
};
#endif
diff --git a/net/lnet/lnet/api-ni.c b/net/lnet/lnet/api-ni.c
index 9d68434..418d65e 100644
--- a/net/lnet/lnet/api-ni.c
+++ b/net/lnet/lnet/api-ni.c
@@ -309,7 +309,7 @@ static int lnet_discover(struct lnet_process_id id, u32 force,
spin_lock_init(&the_lnet.ln_eq_wait_lock);
spin_lock_init(&the_lnet.ln_msg_resend_lock);
init_waitqueue_head(&the_lnet.ln_eq_waitq);
- init_waitqueue_head(&the_lnet.ln_rc_waitq);
+ init_waitqueue_head(&the_lnet.ln_mt_waitq);
mutex_init(&the_lnet.ln_lnd_mutex);
}
@@ -2281,13 +2281,13 @@ void lnet_lib_exit(void)
lnet_ping_target_update(pbuf, ping_mdh);
- rc = lnet_router_checker_start();
+ rc = lnet_monitor_thr_start();
if (rc)
goto err_stop_ping;
rc = lnet_push_target_init();
if (rc != 0)
- goto err_stop_router_checker;
+ goto err_stop_monitor_thr;
rc = lnet_peer_discovery_start();
if (rc != 0)
@@ -2302,8 +2302,8 @@ void lnet_lib_exit(void)
err_destroy_push_target:
lnet_push_target_fini();
-err_stop_router_checker:
- lnet_router_checker_stop();
+err_stop_monitor_thr:
+ lnet_monitor_thr_stop();
err_stop_ping:
lnet_ping_target_fini();
err_acceptor_stop:
@@ -2353,7 +2353,7 @@ void lnet_lib_exit(void)
lnet_router_debugfs_fini();
lnet_peer_discovery_stop();
lnet_push_target_fini();
- lnet_router_checker_stop();
+ lnet_monitor_thr_stop();
lnet_ping_target_fini();
/* Teardown fns that use my own API functions BEFORE here */
diff --git a/net/lnet/lnet/lib-move.c b/net/lnet/lnet/lib-move.c
index 38815fd..418e3ad 100644
--- a/net/lnet/lnet/lib-move.c
+++ b/net/lnet/lnet/lib-move.c
@@ -818,6 +818,9 @@ void lnet_usr_translate_stats(struct lnet_ioctl_element_msg_stats *msg_stats,
}
}
+ /* unset the tx_delay flag as we're going to send it now */
+ msg->msg_tx_delayed = 0;
+
if (do_send) {
lnet_net_unlock(cpt);
lnet_ni_send(ni, msg);
@@ -914,6 +917,9 @@ void lnet_usr_translate_stats(struct lnet_ioctl_element_msg_stats *msg_stats,
msg->msg_niov = rbp->rbp_npages;
msg->msg_kiov = &rb->rb_kiov[0];
+ /* unset the msg-rx_delayed flag since we're receiving the message */
+ msg->msg_rx_delayed = 0;
+
if (do_recv) {
int cpt = msg->msg_rx_cpt;
@@ -2383,6 +2389,98 @@ struct lnet_ni *
return 0;
}
+static int
+lnet_monitor_thread(void *arg)
+{
+ /* The monitor thread takes care of the following:
+ * 1. Checks the aliveness of routers
+ * 2. Checks if there are messages on the resend queue to resend
+ * them.
+ * 3. Check if there are any NIs on the local recovery queue and
+ * pings them
+ * 4. Checks if there are any NIs on the remote recovery queue
+ * and pings them.
+ */
+ while (the_lnet.ln_mt_state == LNET_MT_STATE_RUNNING) {
+ if (lnet_router_checker_active())
+ lnet_check_routers();
+
+ /* TODO do we need to check if we should sleep without
+ * timeout? Technically, an active system will always
+ * have messages in flight so this check will always
+ * evaluate to false. And on an idle system do we care
+ * if we wake up every 1 second? Although, we've seen
+ * cases where we get a complaint that an idle thread
+ * is waking up unnecessarily.
+ */
+ wait_event_interruptible_timeout(the_lnet.ln_mt_waitq,
+ false, HZ);
+ }
+
+ /* clean up the router checker */
+ lnet_prune_rc_data(1);
+
+ /* Shutting down */
+ the_lnet.ln_mt_state = LNET_MT_STATE_SHUTDOWN;
+
+ /* signal that the monitor thread is exiting */
+ complete(&the_lnet.ln_mt_signal);
+
+ return 0;
+}
+
+int lnet_monitor_thr_start(void)
+{
+ int rc;
+ struct task_struct *task;
+
+ LASSERT(the_lnet.ln_mt_state == LNET_MT_STATE_SHUTDOWN);
+
+ init_completion(&the_lnet.ln_mt_signal);
+
+ /* Pre monitor thread start processing */
+ rc = lnet_router_pre_mt_start();
+ if (!rc)
+ return rc;
+
+ the_lnet.ln_mt_state = LNET_MT_STATE_RUNNING;
+ task = kthread_run(lnet_monitor_thread, NULL, "monitor_thread");
+ if (IS_ERR(task)) {
+ rc = PTR_ERR(task);
+ CERROR("Can't start monitor thread: %d\n", rc);
+ /* block until event callback signals exit */
+ wait_for_completion(&the_lnet.ln_mt_signal);
+
+ /* clean up */
+ lnet_router_cleanup();
+ the_lnet.ln_mt_state = LNET_MT_STATE_SHUTDOWN;
+ return -ENOMEM;
+ }
+
+ /* post monitor thread start processing */
+ lnet_router_post_mt_start();
+
+ return 0;
+}
+
+void lnet_monitor_thr_stop(void)
+{
+ if (the_lnet.ln_mt_state == LNET_MT_STATE_SHUTDOWN)
+ return;
+
+ LASSERT(the_lnet.ln_mt_state == LNET_MT_STATE_RUNNING);
+ the_lnet.ln_mt_state = LNET_MT_STATE_STOPPING;
+
+ /* tell the monitor thread that we're shutting down */
+ wake_up(&the_lnet.ln_mt_waitq);
+
+ /* block until monitor thread signals that it's done */
+ wait_for_completion(&the_lnet.ln_mt_signal);
+ LASSERT(the_lnet.ln_mt_state == LNET_MT_STATE_SHUTDOWN);
+
+ lnet_router_cleanup();
+}
+
void
lnet_drop_message(struct lnet_ni *ni, int cpt, void *private, unsigned int nob,
u32 msg_type)
diff --git a/net/lnet/lnet/lib-msg.c b/net/lnet/lnet/lib-msg.c
index a7062f6..7869b96 100644
--- a/net/lnet/lnet/lib-msg.c
+++ b/net/lnet/lnet/lib-msg.c
@@ -141,13 +141,17 @@
{
struct lnet_msg_container *container = the_lnet.ln_msg_containers[cpt];
struct lnet_counters *counters = the_lnet.ln_counters[cpt];
+ s64 timeout_ns;
+
+ /* set the message deadline */
+ timeout_ns = lnet_transaction_timeout * NSEC_PER_SEC;
+ msg->msg_deadline = ktime_add_ns(ktime_get(), timeout_ns);
/* routed message can be committed for both receiving and sending */
LASSERT(!msg->msg_tx_committed);
if (msg->msg_sending) {
LASSERT(!msg->msg_receiving);
-
msg->msg_tx_cpt = cpt;
msg->msg_tx_committed = 1;
if (msg->msg_rx_committed) { /* routed message REPLY */
@@ -161,8 +165,9 @@
}
LASSERT(!msg->msg_onactivelist);
+
msg->msg_onactivelist = 1;
- list_add(&msg->msg_activelist, &container->msc_active);
+ list_add_tail(&msg->msg_activelist, &container->msc_active);
counters->msgs_alloc++;
if (counters->msgs_alloc > counters->msgs_max)
diff --git a/net/lnet/lnet/router.c b/net/lnet/lnet/router.c
index 278807d..3f9d8c5 100644
--- a/net/lnet/lnet/router.c
+++ b/net/lnet/lnet/router.c
@@ -70,9 +70,6 @@
return net->net_tunables.lct_peer_tx_credits;
}
-/* forward ref's */
-static int lnet_router_checker(void *);
-
static int check_routers_before_use;
module_param(check_routers_before_use, int, 0444);
MODULE_PARM_DESC(check_routers_before_use, "Assume routers are down and ping them before use");
@@ -423,8 +420,8 @@ static void lnet_shuffle_seed(void)
if (rnet != rnet2)
kfree(rnet);
- /* indicate to startup the router checker if configured */
- wake_up(&the_lnet.ln_rc_waitq);
+ /* kick start the monitor thread to handle the added route */
+ wake_up(&the_lnet.ln_mt_waitq);
return rc;
}
@@ -809,7 +806,7 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
struct lnet_peer_ni *rtr;
int all_known;
- LASSERT(the_lnet.ln_rc_state == LNET_RC_STATE_RUNNING);
+ LASSERT(the_lnet.ln_mt_state == LNET_MT_STATE_RUNNING);
for (;;) {
int cpt = lnet_net_lock_current();
@@ -1038,7 +1035,7 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
lnet_ni_notify_locked(ni, rtr);
if (!lnet_isrouter(rtr) ||
- the_lnet.ln_rc_state != LNET_RC_STATE_RUNNING) {
+ the_lnet.ln_mt_state != LNET_MT_STATE_RUNNING) {
/* router table changed or router checker is shutting down */
lnet_peer_ni_decref_locked(rtr);
return;
@@ -1092,14 +1089,9 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
lnet_peer_ni_decref_locked(rtr);
}
-int
-lnet_router_checker_start(void)
+int lnet_router_pre_mt_start(void)
{
- struct task_struct *task;
int rc;
- int eqsz = 0;
-
- LASSERT(the_lnet.ln_rc_state == LNET_RC_STATE_SHUTDOWN);
if (check_routers_before_use &&
dead_router_check_interval <= 0) {
@@ -1107,27 +1099,17 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
return -EINVAL;
}
- init_completion(&the_lnet.ln_rc_signal);
-
rc = LNetEQAlloc(0, lnet_router_checker_event, &the_lnet.ln_rc_eqh);
if (rc) {
- CERROR("Can't allocate EQ(%d): %d\n", eqsz, rc);
+ CERROR("Can't allocate EQ(0): %d\n", rc);
return -ENOMEM;
}
- the_lnet.ln_rc_state = LNET_RC_STATE_RUNNING;
- task = kthread_run(lnet_router_checker, NULL, "router_checker");
- if (IS_ERR(task)) {
- rc = PTR_ERR(task);
- CERROR("Can't start router checker thread: %d\n", rc);
- /* block until event callback signals exit */
- wait_for_completion(&the_lnet.ln_rc_signal);
- rc = LNetEQFree(the_lnet.ln_rc_eqh);
- LASSERT(!rc);
- the_lnet.ln_rc_state = LNET_RC_STATE_SHUTDOWN;
- return -ENOMEM;
- }
+ return 0;
+}
+void lnet_router_post_mt_start(void)
+{
if (check_routers_before_use) {
/*
* Note that a helpful side-effect of pinging all known routers
@@ -1136,33 +1118,17 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
*/
lnet_wait_known_routerstate();
}
-
- return 0;
}
-void
-lnet_router_checker_stop(void)
+void lnet_router_cleanup(void)
{
int rc;
- if (the_lnet.ln_rc_state == LNET_RC_STATE_SHUTDOWN)
- return;
-
- LASSERT(the_lnet.ln_rc_state == LNET_RC_STATE_RUNNING);
- the_lnet.ln_rc_state = LNET_RC_STATE_STOPPING;
- /* wakeup the RC thread if it's sleeping */
- wake_up(&the_lnet.ln_rc_waitq);
-
- /* block until event callback signals exit */
- wait_for_completion(&the_lnet.ln_rc_signal);
- LASSERT(the_lnet.ln_rc_state == LNET_RC_STATE_SHUTDOWN);
-
rc = LNetEQFree(the_lnet.ln_rc_eqh);
- LASSERT(!rc);
+ LASSERT(rc == 0);
}
-static void
-lnet_prune_rc_data(int wait_unlink)
+void lnet_prune_rc_data(int wait_unlink)
{
struct lnet_rc_data *rcd;
struct lnet_rc_data *tmp;
@@ -1170,7 +1136,7 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
struct list_head head;
int i = 2;
- if (likely(the_lnet.ln_rc_state == LNET_RC_STATE_RUNNING &&
+ if (likely(the_lnet.ln_mt_state == LNET_MT_STATE_RUNNING &&
list_empty(&the_lnet.ln_rcd_deathrow) &&
list_empty(&the_lnet.ln_rcd_zombie)))
return;
@@ -1179,7 +1145,7 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
lnet_net_lock(LNET_LOCK_EX);
- if (the_lnet.ln_rc_state != LNET_RC_STATE_RUNNING) {
+ if (the_lnet.ln_mt_state != LNET_MT_STATE_RUNNING) {
/* router checker is stopping, prune all */
list_for_each_entry(lp, &the_lnet.ln_routers,
lpni_rtr_list) {
@@ -1242,18 +1208,12 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
}
/*
- * This function is called to check if the RC should block indefinitely.
- * It's called from lnet_router_checker() as well as being passed to
- * wait_event_interruptible() to avoid the lost wake_up problem.
- *
- * When it's called from wait_event_interruptible() it is necessary to
- * also not sleep if the rc state is not running to avoid a deadlock
- * when the system is shutting down
+ * This function is called from the monitor thread to check if there are
+ * any active routers that need to be checked.
*/
-static inline bool
-lnet_router_checker_active(void)
+bool lnet_router_checker_active(void)
{
- if (the_lnet.ln_rc_state != LNET_RC_STATE_RUNNING)
+ if (the_lnet.ln_mt_state != LNET_MT_STATE_RUNNING)
return true;
/*
@@ -1263,70 +1223,54 @@ int lnet_get_rtr_pool_cfg(int idx, struct lnet_ioctl_pool_cfg *pool_cfg)
if (the_lnet.ln_routing)
return true;
+ /* if there are routers that need to be cleaned up then do so */
+ if (!list_empty(&the_lnet.ln_rcd_deathrow) ||
+ !list_empty(&the_lnet.ln_rcd_zombie))
+ return true;
+
return !list_empty(&the_lnet.ln_routers) &&
(live_router_check_interval > 0 ||
dead_router_check_interval > 0);
}
-static int
-lnet_router_checker(void *arg)
+void
+lnet_check_routers(void)
{
struct lnet_peer_ni *rtr;
+ u64 version;
+ int cpt;
+ int cpt2;
- while (the_lnet.ln_rc_state == LNET_RC_STATE_RUNNING) {
- u64 version;
- int cpt;
- int cpt2;
-
- cpt = lnet_net_lock_current();
+ cpt = lnet_net_lock_current();
rescan:
- version = the_lnet.ln_routers_version;
+ version = the_lnet.ln_routers_version;
- list_for_each_entry(rtr, &the_lnet.ln_routers, lpni_rtr_list) {
- cpt2 = rtr->lpni_cpt;
- if (cpt != cpt2) {
- lnet_net_unlock(cpt);
- cpt = cpt2;
- lnet_net_lock(cpt);
- /* the routers list has changed */
- if (version != the_lnet.ln_routers_version)
- goto rescan;
- }
-
- lnet_ping_router_locked(rtr);
-
- /* NB dropped lock */
- if (version != the_lnet.ln_routers_version) {
- /* the routers list has changed */
+ list_for_each_entry(rtr, &the_lnet.ln_routers, lpni_rtr_list) {
+ cpt2 = rtr->lpni_cpt;
+ if (cpt != cpt2) {
+ lnet_net_unlock(cpt);
+ cpt = cpt2;
+ lnet_net_lock(cpt);
+ /* the routers list has changed */
+ if (version != the_lnet.ln_routers_version)
goto rescan;
- }
}
- if (the_lnet.ln_routing)
- lnet_update_ni_status_locked();
-
- lnet_net_unlock(cpt);
-
- lnet_prune_rc_data(0); /* don't wait for UNLINK */
+ lnet_ping_router_locked(rtr);
- /*
- * if there are any routes then wakeup every second. If
- * there are no routes then sleep indefinitely until woken
- * up by a user adding a route
- */
- if (!lnet_router_checker_active())
- wait_event_idle(the_lnet.ln_rc_waitq,
- lnet_router_checker_active());
- else
- schedule_timeout_idle(HZ);
+ /* NB dropped lock */
+ if (version != the_lnet.ln_routers_version) {
+ /* the routers list has changed */
+ goto rescan;
+ }
}
- lnet_prune_rc_data(1); /* wait for UNLINK */
+ if (the_lnet.ln_routing)
+ lnet_update_ni_status_locked();
- the_lnet.ln_rc_state = LNET_RC_STATE_SHUTDOWN;
- complete(&the_lnet.ln_rc_signal);
- /* The unlink event callback will signal final completion */
- return 0;
+ lnet_net_unlock(cpt);
+
+ lnet_prune_rc_data(0); /* don't wait for UNLINK */
}
void
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:31 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:31 -0500
Subject: [lustre-devel] [PATCH 043/622] lustre: checksum: enable/disable
checksum correctly
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-44-git-send-email-jsimmons@infradead.org>
From: Emoly Liu
There are three ways to set checksum support in Lustre. Their
order during client mount is:
- 1. configure --enable/disable-checksum, this(ENABLE_CHECKSUM)
only affects the default mount option and is set in function
client_obd_setup().
- 2. lctl set_param -P osc.*.checksums=0/1, when processing llog,
this value will be set by osc_checksum_seq_write().
- 3. mount option checksum/nochecksum, this will be checked in
ll_options() and be set in client_common_fill_super()->
obd_set_info_async().
This patch fixes one issue in 3. That is if mount option
"-o checksum/nochecksum" is specified, checksum will be changed
accordingly, no matter what is set by "set_param -P" or the
default option; and if no mount option is specified, the value
set by "set_param -P" will be kept. Also, test_77k is added to
sanity.sh to verify this patch.
What's more, a minor initialization issue of cl_supp_cksum_types
is fixed. cl_supp_cksum_types should be always initialized no
matter checksum is enabled or not.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10906
Lustre-commit: e9b13cd1daf9 ("LU-10906 checksum: enable/disable checksum correctly")
Signed-off-by: Emoly Liu
Reviewed-on: https://review.whamcloud.com/32095
Reviewed-by: Yingjin Qian
Reviewed-by: Andreas Dilger
Signed-off-by: James Simmons
---
fs/lustre/ldlm/ldlm_lib.c | 5 +++--
fs/lustre/llite/llite_internal.h | 3 ++-
fs/lustre/llite/llite_lib.c | 23 ++++++++++++++---------
3 files changed, 19 insertions(+), 12 deletions(-)
diff --git a/fs/lustre/ldlm/ldlm_lib.c b/fs/lustre/ldlm/ldlm_lib.c
index 7bc1d10..2c0fad3 100644
--- a/fs/lustre/ldlm/ldlm_lib.c
+++ b/fs/lustre/ldlm/ldlm_lib.c
@@ -355,6 +355,8 @@ int client_obd_setup(struct obd_device *obddev, struct lustre_cfg *lcfg)
init_waitqueue_head(&cli->cl_destroy_waitq);
atomic_set(&cli->cl_destroy_in_flight, 0);
+
+ cli->cl_supp_cksum_types = OBD_CKSUM_CRC32;
/* Turn on checksumming by default. */
cli->cl_checksum = 1;
/*
@@ -362,8 +364,7 @@ int client_obd_setup(struct obd_device *obddev, struct lustre_cfg *lcfg)
* Set cl_chksum* to CRC32 for now to avoid returning screwed info
* through procfs.
*/
- cli->cl_cksum_type = OBD_CKSUM_CRC32;
- cli->cl_supp_cksum_types = OBD_CKSUM_CRC32;
+ cli->cl_cksum_type = cli->cl_supp_cksum_types;
atomic_set(&cli->cl_resends, OSC_DEFAULT_RESENDS);
/*
diff --git a/fs/lustre/llite/llite_internal.h b/fs/lustre/llite/llite_internal.h
index d0a703d..6bdbf28 100644
--- a/fs/lustre/llite/llite_internal.h
+++ b/fs/lustre/llite/llite_internal.h
@@ -479,7 +479,8 @@ struct ll_sb_info {
unsigned int ll_umounting:1,
ll_xattr_cache_enabled:1,
ll_xattr_cache_set:1, /* already set to 0/1 */
- ll_client_common_fill_super_succeeded:1;
+ ll_client_common_fill_super_succeeded:1,
+ ll_checksum_set:1;
struct lustre_client_ocd ll_lco;
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index e2c7a4d..eb29064 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -560,13 +560,15 @@ static int client_common_fill_super(struct super_block *sb, char *md, char *dt)
}
checksum = sbi->ll_flags & LL_SBI_CHECKSUM;
- err = obd_set_info_async(NULL, sbi->ll_dt_exp, sizeof(KEY_CHECKSUM),
- KEY_CHECKSUM, sizeof(checksum), &checksum,
- NULL);
- if (err) {
- CERROR("%s: Set checksum failed: rc = %d\n",
- sbi->ll_dt_exp->exp_obd->obd_name, err);
- goto out_root;
+ if (sbi->ll_checksum_set) {
+ err = obd_set_info_async(NULL, sbi->ll_dt_exp,
+ sizeof(KEY_CHECKSUM), KEY_CHECKSUM,
+ sizeof(checksum), &checksum, NULL);
+ if (err) {
+ CERROR("%s: Set checksum failed: rc = %d\n",
+ sbi->ll_dt_exp->exp_obd->obd_name, err);
+ goto out_root;
+ }
}
cl_sb_init(sb);
@@ -763,10 +765,11 @@ static inline int ll_set_opt(const char *opt, char *data, int fl)
}
/* non-client-specific mount options are parsed in lmd_parse */
-static int ll_options(char *options, int *flags)
+static int ll_options(char *options, struct ll_sb_info *sbi)
{
int tmp;
char *s1 = options, *s2;
+ int *flags = &sbi->ll_flags;
if (!options)
return 0;
@@ -832,11 +835,13 @@ static int ll_options(char *options, int *flags)
tmp = ll_set_opt("checksum", s1, LL_SBI_CHECKSUM);
if (tmp) {
*flags |= tmp;
+ sbi->ll_checksum_set = 1;
goto next;
}
tmp = ll_set_opt("nochecksum", s1, LL_SBI_CHECKSUM);
if (tmp) {
*flags &= ~tmp;
+ sbi->ll_checksum_set = 1;
goto next;
}
tmp = ll_set_opt("lruresize", s1, LL_SBI_LRU_RESIZE);
@@ -971,7 +976,7 @@ int ll_fill_super(struct super_block *sb)
goto out_free;
}
- err = ll_options(lsi->lsi_lmd->lmd_opts, &sbi->ll_flags);
+ err = ll_options(lsi->lsi_lmd->lmd_opts, sbi);
if (err)
goto out_free;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:41 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:41 -0500
Subject: [lustre-devel] [PATCH 053/622] lustre: osc: truncate does not
update blocks count on client
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-54-git-send-email-jsimmons@infradead.org>
From: Arshad Hussain
'truncate' call correctly updates the server side with
correct size and blocks count. However, on the client
side all the metadata are correctly updated except the
blocks count, which still reflects the old count prior
to truncate call. This patch fixes this issue on the
client by modifying osc_io_setattr_end() to update
attr with the updated block count.
New test case under sanity is added to verify the that
the blocks counts are correctly updated after truncate call
Co-authored-by: Abrarahmed Momin
WC-bug-id: https://jira.whamcloud.com/browse/LU-10370
Lustre-commit: 6115eb7fd55a ("LU-10370 ofd: truncate does not update blocks count on client")
Signed-off-by: Abrarahmed Momin
Signed-off-by: Arshad Hussain
Reviewed-on: https://review.whamcloud.com/31073
Reviewed-by: Jinshan Xiong
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/osc/osc_io.c | 11 +++++++++++
1 file changed, 11 insertions(+)
diff --git a/fs/lustre/osc/osc_io.c b/fs/lustre/osc/osc_io.c
index 970e8a7..1485962 100644
--- a/fs/lustre/osc/osc_io.c
+++ b/fs/lustre/osc/osc_io.c
@@ -588,6 +588,9 @@ void osc_io_setattr_end(const struct lu_env *env,
struct osc_io *oio = cl2osc_io(env, slice);
struct cl_object *obj = slice->cis_obj;
struct osc_async_cbargs *cbargs = &oio->oi_cbarg;
+ struct cl_attr *attr = &osc_env_info(env)->oti_attr;
+ struct obdo *oa = &oio->oi_oa;
+ unsigned int cl_valid = 0;
int result = 0;
if (cbargs->opc_rpc_sent) {
@@ -609,6 +612,14 @@ void osc_io_setattr_end(const struct lu_env *env,
if (cl_io_is_trunc(io)) {
u64 size = io->u.ci_setattr.sa_attr.lvb_size;
+ cl_object_attr_lock(obj);
+ if (oa->o_valid & OBD_MD_FLBLOCKS) {
+ attr->cat_blocks = oa->o_blocks;
+ cl_valid |= CAT_BLOCKS;
+ }
+
+ cl_object_attr_update(env, obj, attr, cl_valid);
+ cl_object_attr_unlock(obj);
osc_trunc_check(env, io, oio, size);
osc_cache_truncate_end(env, oio->oi_trunc);
oio->oi_trunc = NULL;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:49 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:49 -0500
Subject: [lustre-devel] [PATCH 001/622] lustre: always enable special
debugging, fhandles, and quota support.
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-2-git-send-email-jsimmons@infradead.org>
Lustre heavily depends on fhandles for its FID handling and needs
quota always enabled.
Signed-off-by: James Simmons
---
fs/lustre/Kconfig | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/fs/lustre/Kconfig b/fs/lustre/Kconfig
index 2ea3f24..2eb7e45 100644
--- a/fs/lustre/Kconfig
+++ b/fs/lustre/Kconfig
@@ -9,6 +9,9 @@ config LUSTRE_FS
select CRYPTO_SHA1
select CRYPTO_SHA256
select CRYPTO_SHA512
+ select DEBUG_FS
+ select FHANDLE
+ select QUOTA
depends on MULTIUSER
help
This option enables Lustre file system client support. Choose Y
@@ -43,6 +46,7 @@ config LUSTRE_FS_POSIX_ACL
config LUSTRE_DEBUG_EXPENSIVE_CHECK
bool "Enable Lustre DEBUG checks"
+ select REFCOUNT_FULL
depends on LUSTRE_FS
help
This option is mainly for debug purpose. It enables Lustre code to do
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:47 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:47 -0500
Subject: [lustre-devel] [PATCH 059/622] lustre: ptlrpc: don't zero request
handle
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-60-git-send-email-jsimmons@infradead.org>
From: Alexander Boyko
LNet can retransmit a request at any time if it isn't replied.
The ptlrpc_resend_req zero the request handle and ptlrpc_send_rpc
set it. If retransmission happen with zeroed handle, the client
can't find a valid export by handle and set rq_export to NULL and
reply with ENOTCONN. A server evict client with this error.
client (nid x.x.x.x at tcp) returned error from blocking AST
(req status -107 rc -107), evict it
WC-bug-id: https://jira.whamcloud.com/browse/LU-11117
Lustre-commit: 00c72ab6bb43 ("LU-11117 ptlrpc: don't zero request handle")
Signed-off-by: Alexander Boyko
Cray-bug-id: LUS-6037
Reviewed-on: https://review.whamcloud.com/32781
Reviewed-by: Mikhail Pershin
Reviewed-by: Alexey Lyashkov
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/ptlrpc/client.c | 1 -
1 file changed, 1 deletion(-)
diff --git a/fs/lustre/ptlrpc/client.c b/fs/lustre/ptlrpc/client.c
index 9b41c12..d28a9cd 100644
--- a/fs/lustre/ptlrpc/client.c
+++ b/fs/lustre/ptlrpc/client.c
@@ -2728,7 +2728,6 @@ void ptlrpc_resend_req(struct ptlrpc_request *req)
return;
}
- lustre_msg_set_handle(req->rq_reqmsg, &(struct lustre_handle){ 0 });
req->rq_status = -EAGAIN;
req->rq_resend = 1;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:50 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:50 -0500
Subject: [lustre-devel] [PATCH 062/622] lustre: ptlrpc: remove obsolete OBD
RPC opcodes
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-63-git-send-email-jsimmons@infradead.org>
From: Andreas Dilger
Remove the obsolete OBD_LOG_CANCEL (since Lustre 1.5) and
OBD_QC_CALLBACK (since Lustre 2.4) RPC opcodes.
Assign OBD_IDX_READ an explicit opcode (as should be done with all
enums in lustre_idl.h) so that the value does not change if some
prior field is removed.
Also remove the OBD_FAIL checks that were used to test them.
The setting in conf_sanity.sh test_58 was unused for many years.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10855
Lustre-commit: 7d89a5b8aefc ("LU-10855 ptlrpc: remove obsolete OBD RPC opcodes")
Signed-off-by: Andreas Dilger
Reviewed-on: https://review.whamcloud.com/32651
Reviewed-by: John L. Hammond
Reviewed-by: James Simmons
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/include/obd_support.h | 6 +++---
fs/lustre/ptlrpc/lproc_ptlrpc.c | 4 ++--
fs/lustre/ptlrpc/wiretest.c | 4 ----
include/uapi/linux/lustre/lustre_idl.h | 12 ++++++------
4 files changed, 11 insertions(+), 15 deletions(-)
diff --git a/fs/lustre/include/obd_support.h b/fs/lustre/include/obd_support.h
index 67500b5..99b4f1f 100644
--- a/fs/lustre/include/obd_support.h
+++ b/fs/lustre/include/obd_support.h
@@ -352,12 +352,12 @@
#define OBD_FAIL_PTLRPC_BULK_ATTACH 0x521
#define OBD_FAIL_OBD_PING_NET 0x600
-#define OBD_FAIL_OBD_LOG_CANCEL_NET 0x601
+/* OBD_FAIL_OBD_LOG_CANCEL_NET 0x601 obsolete since 1.5 */
#define OBD_FAIL_OBD_LOGD_NET 0x602
-/* OBD_FAIL_OBD_QC_CALLBACK_NET 0x603 obsolete since 2.4 */
+/* OBD_FAIL_OBD_QC_CALLBACK_NET 0x603 obsolete since 2.4 */
#define OBD_FAIL_OBD_DQACQ 0x604
#define OBD_FAIL_OBD_LLOG_SETUP 0x605
-#define OBD_FAIL_OBD_LOG_CANCEL_REP 0x606
+/* OBD_FAIL_OBD_LOG_CANCEL_REP 0x606 obsolete since 1.5 */
#define OBD_FAIL_OBD_IDX_READ_NET 0x607
#define OBD_FAIL_OBD_IDX_READ_BREAK 0x608
#define OBD_FAIL_OBD_NO_LRU 0x609
diff --git a/fs/lustre/ptlrpc/lproc_ptlrpc.c b/fs/lustre/ptlrpc/lproc_ptlrpc.c
index 0efbcfc..b70a1c7 100644
--- a/fs/lustre/ptlrpc/lproc_ptlrpc.c
+++ b/fs/lustre/ptlrpc/lproc_ptlrpc.c
@@ -111,8 +111,8 @@
{ MGS_SET_INFO, "mgs_set_info" },
{ MGS_CONFIG_READ, "mgs_config_read" },
{ OBD_PING, "obd_ping" },
- { OBD_LOG_CANCEL, "llog_cancel" },
- { OBD_QC_CALLBACK, "obd_quota_callback" },
+ { 401, /* was OBD_LOG_CANCEL */ "llog_cancel" },
+ { 402, /* was OBD_QC_CALLBACK */ "obd_quota_callback" },
{ OBD_IDX_READ, "dt_index_read" },
{ LLOG_ORIGIN_HANDLE_CREATE, "llog_origin_handle_open" },
{ LLOG_ORIGIN_HANDLE_NEXT_BLOCK, "llog_origin_handle_next_block" },
diff --git a/fs/lustre/ptlrpc/wiretest.c b/fs/lustre/ptlrpc/wiretest.c
index 202c5ab..015c5bd 100644
--- a/fs/lustre/ptlrpc/wiretest.c
+++ b/fs/lustre/ptlrpc/wiretest.c
@@ -326,10 +326,6 @@ void lustre_assert_wire_constants(void)
BUILD_BUG_ON(LUSTRE_RES_ID_HSH_OFF != 3);
LASSERTF(OBD_PING == 400, "found %lld\n",
(long long)OBD_PING);
- LASSERTF(OBD_LOG_CANCEL == 401, "found %lld\n",
- (long long)OBD_LOG_CANCEL);
- LASSERTF(OBD_QC_CALLBACK == 402, "found %lld\n",
- (long long)OBD_QC_CALLBACK);
LASSERTF(OBD_IDX_READ == 403, "found %lld\n",
(long long)OBD_IDX_READ);
LASSERTF(OBD_LAST_OPC == 404, "found %lld\n",
diff --git a/include/uapi/linux/lustre/lustre_idl.h b/include/uapi/linux/lustre/lustre_idl.h
index 798aa57..adaa994 100644
--- a/include/uapi/linux/lustre/lustre_idl.h
+++ b/include/uapi/linux/lustre/lustre_idl.h
@@ -2342,13 +2342,13 @@ struct cfg_marker {
*/
enum obd_cmd {
- OBD_PING = 400,
- OBD_LOG_CANCEL, /* Obsolete since 1.5. */
- OBD_QC_CALLBACK, /* not used since 2.4 */
- OBD_IDX_READ,
- OBD_LAST_OPC
+ OBD_PING = 400,
+/* OBD_LOG_CANCEL = 401, Obsolete since 1.5 */
+/* OBD_QC_CALLBACK = 402, not used since 2.4 */
+ OBD_IDX_READ = 403,
+ OBD_LAST_OPC,
+ OBD_FIRST_OPC = OBD_PING
};
-#define OBD_FIRST_OPC OBD_PING
/**
* llog contexts indices.
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:53 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:53 -0500
Subject: [lustre-devel] [PATCH 065/622] lustre: osc: fix idle_timeout
handling
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-66-git-send-email-jsimmons@infradead.org>
The patch that landed for LU-7236 introduced new sysfs entries
which were done wrong.
1) For idle_timeout it returns -ERANGE for
any value passed in expect setting idle_timeout to zero. This
does not match what the commit message said for LU-7236. So
I changed lprocfs_str_with_units_to_s64() into kstrtouint()
since a signed 64 bit timeout is not needed. Using kstrtouint()
ensures that negative values are not possible and also cap the
value to CONNECTION_SWITCH_MAX since the max of 4 billion
seconds is over kill.
2) For the next procfs idle_connect it is really a write only file
but it was treated as both read and write. There is no need for
the osc_idle_connect_seq_show() function.
3) Lastly no more stuffing new entries into proc or debugfs. For
this patch convert these new proc entries to sysfs. It seems
to be a common occurrence so add LPROC_SEQ_* to spelling.txt
so checkpatch will complain about using LPROC_SEQ_* which will
go away.
WC-bug-id: https://jira.whamcloud.com/browse/LU-8066
Lustre-commit: 406cd8a74d84 ("LU-8066 osc: fix idle_timeout handling")
Signed-off-by: James Simmons
Reviewed-on: https://review.whamcloud.com/32719
Reviewed-by: Alex Zhuravlev
Reviewed-by: John L. Hammond
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/osc/lproc_osc.c | 42 ++++++++++++++++++------------------------
1 file changed, 18 insertions(+), 24 deletions(-)
diff --git a/fs/lustre/osc/lproc_osc.c b/fs/lustre/osc/lproc_osc.c
index fd84393..0a12079 100644
--- a/fs/lustre/osc/lproc_osc.c
+++ b/fs/lustre/osc/lproc_osc.c
@@ -598,26 +598,27 @@ static int osc_unstable_stats_seq_show(struct seq_file *m, void *v)
LPROC_SEQ_FOPS_RO(osc_unstable_stats);
-static int osc_idle_timeout_seq_show(struct seq_file *m, void *v)
+static ssize_t idle_timeout_show(struct kobject *kobj, struct attribute *attr,
+ char *buf)
{
- struct obd_device *obd = m->private;
+ struct obd_device *obd = container_of(kobj, struct obd_device,
+ obd_kset.kobj);
struct client_obd *cli = &obd->u.cli;
- seq_printf(m, "%u\n", cli->cl_import->imp_idle_timeout);
- return 0;
+ return sprintf(buf, "%u\n", cli->cl_import->imp_idle_timeout);
}
-static ssize_t osc_idle_timeout_seq_write(struct file *f,
- const char __user *buffer,
- size_t count, loff_t *off)
+static ssize_t idle_timeout_store(struct kobject *kobj, struct attribute *attr,
+ const char *buffer, size_t count)
{
- struct obd_device *obd = ((struct seq_file *)f->private_data)->private;
+ struct obd_device *obd = container_of(kobj, struct obd_device,
+ obd_kset.kobj);
struct client_obd *cli = &obd->u.cli;
struct ptlrpc_request *req;
unsigned int val;
int rc;
- rc = kstrtouint_from_user(buffer, count, 0, &val);
+ rc = kstrtouint(buffer, 0, &val);
if (rc)
return rc;
@@ -635,18 +636,13 @@ static ssize_t osc_idle_timeout_seq_write(struct file *f,
return count;
}
-LPROC_SEQ_FOPS(osc_idle_timeout);
+LUSTRE_RW_ATTR(idle_timeout);
-static int osc_idle_connect_seq_show(struct seq_file *m, void *v)
+static ssize_t idle_connect_store(struct kobject *kobj, struct attribute *attr,
+ const char *buffer, size_t count)
{
- return 0;
-}
-
-static ssize_t osc_idle_connect_seq_write(struct file *f,
- const char __user *buffer,
- size_t count, loff_t *off)
-{
- struct obd_device *dev = ((struct seq_file *)f->private_data)->private;
+ struct obd_device *dev = container_of(kobj, struct obd_device,
+ obd_kset.kobj);
struct client_obd *cli = &dev->u.cli;
struct ptlrpc_request *req;
@@ -658,7 +654,7 @@ static ssize_t osc_idle_connect_seq_write(struct file *f,
return count;
}
-LPROC_SEQ_FOPS(osc_idle_connect);
+LUSTRE_WO_ATTR(idle_connect);
LPROC_SEQ_FOPS_RO_TYPE(osc, connect_flags);
LPROC_SEQ_FOPS_RO_TYPE(osc, server_uuid);
@@ -687,10 +683,6 @@ static ssize_t osc_idle_connect_seq_write(struct file *f,
.fops = &osc_pinger_recov_fops },
{ .name = "unstable_stats",
.fops = &osc_unstable_stats_fops },
- { .name = "idle_timeout",
- .fops = &osc_idle_timeout_fops },
- { .name = "idle_connect",
- .fops = &osc_idle_connect_fops },
{ NULL }
};
@@ -877,6 +869,8 @@ void lproc_osc_attach_seqstat(struct obd_device *dev)
&lustre_attr_resend_count.attr,
&lustre_attr_ost_conn_uuid.attr,
&lustre_attr_ping.attr,
+ &lustre_attr_idle_timeout.attr,
+ &lustre_attr_idle_connect.attr,
NULL,
};
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:55 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:55 -0500
Subject: [lustre-devel] [PATCH 067/622] lustre: obd: keep dirty_max_pages a
round number of MB
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-68-git-send-email-jsimmons@infradead.org>
From: "John L. Hammond"
In client_adjust_max_dirty() ensure that the dirty pages limit is
always divisible by 256 so that it may faithfully be represented in MB
as is the case when the max_dirty_mb parameters are used.
WC-bug-id: https://jira.whamcloud.com/browse/LU-11157
Lustre-commit: d3f88d376c49 ("LU-11157 obd: keep dirty_max_pages a round number of MB")
Signed-off-by: John L. Hammond
Reviewed-on: https://review.whamcloud.com/32831
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Reviewed-by: James Simmons
Signed-off-by: James Simmons
---
fs/lustre/include/obd.h | 13 ++++++++++---
1 file changed, 10 insertions(+), 3 deletions(-)
diff --git a/fs/lustre/include/obd.h b/fs/lustre/include/obd.h
index d2bd234..5656eb0 100644
--- a/fs/lustre/include/obd.h
+++ b/fs/lustre/include/obd.h
@@ -1106,7 +1106,7 @@ static inline int cli_brw_size(struct obd_device *obd)
}
/*
- * when RPC size or the max RPCs in flight is increased, the max dirty pages
+ * When RPC size or the max RPCs in flight is increased, the max dirty pages
* of the client should be increased accordingly to avoid sending fragmented
* RPCs over the network when the client runs out of the maximum dirty space
* when so many RPCs are being generated.
@@ -1114,10 +1114,10 @@ static inline int cli_brw_size(struct obd_device *obd)
static inline void client_adjust_max_dirty(struct client_obd *cli)
{
/* initializing */
- if (cli->cl_dirty_max_pages <= 0)
+ if (cli->cl_dirty_max_pages <= 0) {
cli->cl_dirty_max_pages =
(OSC_MAX_DIRTY_DEFAULT * 1024 * 1024) >> PAGE_SHIFT;
- else {
+ } else {
unsigned long dirty_max = cli->cl_max_rpcs_in_flight *
cli->cl_max_pages_per_rpc;
@@ -1127,6 +1127,13 @@ static inline void client_adjust_max_dirty(struct client_obd *cli)
if (cli->cl_dirty_max_pages > totalram_pages() / 8)
cli->cl_dirty_max_pages = totalram_pages() / 8;
+
+ /* This value is exported to userspace through the max_dirty_mb
+ * parameter. So we round up the number of pages to make it a round
+ * number of MBs.
+ */
+ cli->cl_dirty_max_pages = round_up(cli->cl_dirty_max_pages,
+ 1 << (20 - PAGE_SHIFT));
}
#endif /* __OBD_H */
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:56 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:56 -0500
Subject: [lustre-devel] [PATCH 068/622] lustre: osc: depart grant shrinking
from pinger
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-69-git-send-email-jsimmons@infradead.org>
From: Bobi Jam
* Removing grant shrinking code outside of pinger, use a workqueue
to handle grant shrinking timer.
* Enable OSC grant shrinking by default.
bugzilla: 19507
WC-bug-id: https://jira.whamcloud.com/browse/LU-8708
Lustre-commit: fc915a43786e ("LU-8708 osc: depart grant shrinking from pinger")
Signed-off-by: Bobi Jam
Reviewed-on: https://review.whamcloud.com/23202
Reviewed-by: Hongchao Zhang
Reviewed-by: Andreas Dilger
Reviewed-by: James Simmons
Signed-off-by: James Simmons
---
fs/lustre/ldlm/ldlm_lib.c | 1 +
fs/lustre/llite/llite_lib.c | 2 +-
fs/lustre/osc/osc_request.c | 155 ++++++++++++++++++++++++++++++--------------
3 files changed, 110 insertions(+), 48 deletions(-)
diff --git a/fs/lustre/ldlm/ldlm_lib.c b/fs/lustre/ldlm/ldlm_lib.c
index 2c0fad3..838ddb3 100644
--- a/fs/lustre/ldlm/ldlm_lib.c
+++ b/fs/lustre/ldlm/ldlm_lib.c
@@ -349,6 +349,7 @@ int client_obd_setup(struct obd_device *obddev, struct lustre_cfg *lcfg)
spin_lock_init(&cli->cl_lru_list_lock);
atomic_long_set(&cli->cl_unstable_count, 0);
INIT_LIST_HEAD(&cli->cl_shrink_list);
+ INIT_LIST_HEAD(&cli->cl_grant_chain);
INIT_LIST_HEAD(&cli->cl_flight_waiters);
cli->cl_rpcs_in_flight = 0;
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index 0844318..56624e8 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -399,7 +399,7 @@ static int client_common_fill_super(struct super_block *sb, char *md, char *dt)
OBD_CONNECT_LAYOUTLOCK |
OBD_CONNECT_PINGLESS | OBD_CONNECT_LFSCK |
OBD_CONNECT_BULK_MBITS | OBD_CONNECT_SHORTIO |
- OBD_CONNECT_FLAGS2;
+ OBD_CONNECT_FLAGS2 | OBD_CONNECT_GRANT_SHRINK;
/* The client currently advertises support for OBD_CONNECT_LOCKAHEAD_OLD
* so it can interoperate with an older version of lockahead which was
diff --git a/fs/lustre/osc/osc_request.c b/fs/lustre/osc/osc_request.c
index e341fcc..1a9ed8d 100644
--- a/fs/lustre/osc/osc_request.c
+++ b/fs/lustre/osc/osc_request.c
@@ -33,6 +33,7 @@
#define DEBUG_SUBSYSTEM S_OSC
+#include
#include
#include
#include
@@ -721,6 +722,16 @@ static void osc_update_grant(struct client_obd *cli, struct ost_body *body)
}
}
+/**
+ * grant thread data for shrinking space.
+ */
+struct grant_thread_data {
+ struct list_head gtd_clients;
+ struct mutex gtd_mutex;
+ unsigned long gtd_stopped:1;
+};
+static struct grant_thread_data client_gtd;
+
static int osc_shrink_grant_interpret(const struct lu_env *env,
struct ptlrpc_request *req,
void *aa, int rc)
@@ -823,6 +834,9 @@ static int osc_should_shrink_grant(struct client_obd *client)
{
time64_t next_shrink = client->cl_next_shrink_grant;
+ if (!client->cl_import)
+ return 0;
+
if ((client->cl_import->imp_connect_data.ocd_connect_flags &
OBD_CONNECT_GRANT_SHRINK) == 0)
return 0;
@@ -843,38 +857,83 @@ static int osc_should_shrink_grant(struct client_obd *client)
return 0;
}
-static int osc_grant_shrink_grant_cb(struct timeout_item *item, void *data)
-{
- struct client_obd *client;
+#define GRANT_SHRINK_RPC_BATCH 100
+
+static void osc_grant_work_handler(struct work_struct *data);
+static DECLARE_DELAYED_WORK(work, osc_grant_work_handler);
- list_for_each_entry(client, &item->ti_obd_list, cl_grant_shrink_list) {
- if (osc_should_shrink_grant(client))
- osc_shrink_grant(client);
+static void osc_grant_work_handler(struct work_struct *data)
+{
+ struct client_obd *cli;
+ int rpc_sent;
+ bool init_next_shrink = true;
+ time64_t next_shrink = ktime_get_seconds() + GRANT_SHRINK_INTERVAL;
+
+ rpc_sent = 0;
+ mutex_lock(&client_gtd.gtd_mutex);
+ list_for_each_entry(cli, &client_gtd.gtd_clients,
+ cl_grant_chain) {
+ if (++rpc_sent < GRANT_SHRINK_RPC_BATCH &&
+ osc_should_shrink_grant(cli))
+ osc_shrink_grant(cli);
+
+ if (!init_next_shrink) {
+ if (cli->cl_next_shrink_grant < next_shrink &&
+ cli->cl_next_shrink_grant > ktime_get_seconds())
+ next_shrink = cli->cl_next_shrink_grant;
+ } else {
+ init_next_shrink = false;
+ next_shrink = cli->cl_next_shrink_grant;
+ }
}
- return 0;
+ mutex_unlock(&client_gtd.gtd_mutex);
+
+ if (client_gtd.gtd_stopped == 1)
+ return;
+
+ if (next_shrink > ktime_get_seconds())
+ schedule_delayed_work(&work, msecs_to_jiffies(
+ (next_shrink - ktime_get_seconds()) *
+ MSEC_PER_SEC));
+ else
+ schedule_work(&work.work);
}
-static int osc_add_shrink_grant(struct client_obd *client)
+/**
+ * Start grant thread for returing grant to server for idle clients.
+ */
+static int osc_start_grant_work(void)
{
- int rc;
+ client_gtd.gtd_stopped = 0;
+ mutex_init(&client_gtd.gtd_mutex);
+ INIT_LIST_HEAD(&client_gtd.gtd_clients);
+
+ schedule_work(&work.work);
- rc = ptlrpc_add_timeout_client(client->cl_grant_shrink_interval,
- TIMEOUT_GRANT,
- osc_grant_shrink_grant_cb, NULL,
- &client->cl_grant_shrink_list);
- if (rc) {
- CERROR("add grant client %s error %d\n", cli_name(client), rc);
- return rc;
- }
- CDEBUG(D_CACHE, "add grant client %s\n", cli_name(client));
- osc_update_next_shrink(client);
return 0;
}
-static int osc_del_shrink_grant(struct client_obd *client)
+static void osc_stop_grant_work(void)
+{
+ client_gtd.gtd_stopped = 1;
+ cancel_delayed_work_sync(&work);
+}
+
+static void osc_add_grant_list(struct client_obd *client)
{
- return ptlrpc_del_timeout_client(&client->cl_grant_shrink_list,
- TIMEOUT_GRANT);
+ mutex_lock(&client_gtd.gtd_mutex);
+ list_add(&client->cl_grant_chain, &client_gtd.gtd_clients);
+ mutex_unlock(&client_gtd.gtd_mutex);
+}
+
+static void osc_del_grant_list(struct client_obd *client)
+{
+ if (list_empty(&client->cl_grant_chain))
+ return;
+
+ mutex_lock(&client_gtd.gtd_mutex);
+ list_del_init(&client->cl_grant_chain);
+ mutex_unlock(&client_gtd.gtd_mutex);
}
void osc_init_grant(struct client_obd *cli, struct obd_connect_data *ocd)
@@ -929,9 +988,8 @@ void osc_init_grant(struct client_obd *cli, struct obd_connect_data *ocd)
cli_name(cli), cli->cl_avail_grant, cli->cl_lost_grant,
cli->cl_chunkbits, cli->cl_max_extent_pages);
- if (ocd->ocd_connect_flags & OBD_CONNECT_GRANT_SHRINK &&
- list_empty(&cli->cl_grant_shrink_list))
- osc_add_shrink_grant(cli);
+ if (OCD_HAS_FLAG(ocd, GRANT_SHRINK) && list_empty(&cli->cl_grant_chain))
+ osc_add_grant_list(cli);
}
EXPORT_SYMBOL(osc_init_grant);
@@ -2971,15 +3029,12 @@ int osc_disconnect(struct obd_export *exp)
* osc_disconnect
* del_shrink_grant
* ptlrpc_connect_interrupt
- * init_grant_shrink
+ * osc_init_grant
* add this client to shrink list
- * cleanup_osc
- * Bang! pinger trigger the shrink.
- * So the osc should be disconnected from the shrink list, after we
- * are sure the import has been destroyed. BUG18662
+ * cleanup_osc
+ * Bang! grant shrink thread trigger the shrink. BUG18662
*/
- if (!obd->u.cli.cl_import)
- osc_del_shrink_grant(&obd->u.cli);
+ osc_del_grant_list(&obd->u.cli);
return rc;
}
EXPORT_SYMBOL(osc_disconnect);
@@ -3159,8 +3214,8 @@ int osc_setup_common(struct obd_device *obd, struct lustre_cfg *lcfg)
goto out_ptlrpcd_work;
cli->cl_grant_shrink_interval = GRANT_SHRINK_INTERVAL;
+ osc_update_next_shrink(cli);
- INIT_LIST_HEAD(&cli->cl_grant_shrink_list);
return 0;
out_ptlrpcd_work:
@@ -3210,7 +3265,6 @@ int osc_setup(struct obd_device *obd, struct lustre_cfg *lcfg)
atomic_add(added, &osc_pool_req_count);
}
- INIT_LIST_HEAD(&cli->cl_grant_shrink_list);
ns_register_cancel(obd->obd_namespace, osc_cancel_weight);
spin_lock(&osc_shrink_lock);
@@ -3356,14 +3410,19 @@ static int __init osc_init(void)
if (rc)
return rc;
+ rc = class_register_type(&osc_obd_ops, NULL,
+ LUSTRE_OSC_NAME, &osc_device_type);
+ if (rc)
+ goto out_kmem;
+
rc = register_shrinker(&osc_cache_shrinker);
if (rc)
- goto err;
+ goto out_type;
/* This is obviously too much memory, only prevent overflow here */
if (osc_reqpool_mem_max >= 1 << 12 || osc_reqpool_mem_max == 0) {
rc = -EINVAL;
- goto err;
+ goto out_shrinker;
}
reqpool_size = osc_reqpool_mem_max << 20;
@@ -3383,29 +3442,31 @@ static int __init osc_init(void)
atomic_set(&osc_pool_req_count, 0);
osc_rq_pool = ptlrpc_init_rq_pool(0, OST_MAXREQSIZE,
ptlrpc_add_rqs_to_pool);
+ if (!osc_rq_pool) {
+ rc = -ENOMEM;
+ goto out_shrinker;
+ }
- rc = -ENOMEM;
-
- if (!osc_rq_pool)
- goto err;
-
- rc = class_register_type(&osc_obd_ops, NULL,
- LUSTRE_OSC_NAME, &osc_device_type);
+ rc = osc_start_grant_work();
if (rc)
- goto err;
+ goto out_req_pool;
return rc;
-err:
- if (osc_rq_pool)
- ptlrpc_free_rq_pool(osc_rq_pool);
+out_req_pool:
+ ptlrpc_free_rq_pool(osc_rq_pool);
+out_type:
+ class_unregister_type(LUSTRE_OSC_NAME);
+out_shrinker:
unregister_shrinker(&osc_cache_shrinker);
+out_kmem:
lu_kmem_fini(osc_caches);
return rc;
}
static void /*__exit*/ osc_exit(void)
{
+ osc_stop_grant_work();
unregister_shrinker(&osc_cache_shrinker);
class_unregister_type(LUSTRE_OSC_NAME);
lu_kmem_fini(osc_caches);
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:09:08 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:09:08 -0500
Subject: [lustre-devel] [PATCH 080/622] lnet: handle o2iblnd tx failure
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-81-git-send-email-jsimmons@infradead.org>
From: Amir Shehata
Monitor the different types of failures that might occur on the
transmit and flag the type of failure to be propagated to LNet
which will handle either by attempting a resend or simply
finalizing the message and propagating a failure to the ULP.
WC-bug-id: https://jira.whamcloud.com/browse/LU-9120
Lustre-commit: 8cf835e425d8 ("LU-9120 lnet: handle o2iblnd tx failure")
Signed-off-by: Amir Shehata
Reviewed-on: https://review.whamcloud.com/32765
Reviewed-by: Sonia Sharma
Reviewed-by: Olaf Weber
Signed-off-by: James Simmons
---
net/lnet/klnds/o2iblnd/o2iblnd.c | 2 +-
net/lnet/klnds/o2iblnd/o2iblnd.h | 4 ++-
net/lnet/klnds/o2iblnd/o2iblnd_cb.c | 59 ++++++++++++++++++++++++++++++++-----
3 files changed, 55 insertions(+), 10 deletions(-)
diff --git a/net/lnet/klnds/o2iblnd/o2iblnd.c b/net/lnet/klnds/o2iblnd/o2iblnd.c
index 825fe30..017fe5f 100644
--- a/net/lnet/klnds/o2iblnd/o2iblnd.c
+++ b/net/lnet/klnds/o2iblnd/o2iblnd.c
@@ -519,7 +519,7 @@ static int kiblnd_del_peer(struct lnet_ni *ni, lnet_nid_t nid)
write_unlock_irqrestore(&kiblnd_data.kib_global_lock, flags);
- kiblnd_txlist_done(&zombies, -EIO);
+ kiblnd_txlist_done(&zombies, -EIO, LNET_MSG_STATUS_LOCAL_ERROR);
return rc;
}
diff --git a/net/lnet/klnds/o2iblnd/o2iblnd.h b/net/lnet/klnds/o2iblnd/o2iblnd.h
index 9021051..999b58d 100644
--- a/net/lnet/klnds/o2iblnd/o2iblnd.h
+++ b/net/lnet/klnds/o2iblnd/o2iblnd.h
@@ -515,6 +515,7 @@ struct kib_tx { /* transmit message */
short tx_queued; /* queued for sending */
short tx_waiting; /* waiting for peer_ni */
int tx_status; /* LNET completion status */
+ enum lnet_msg_hstatus tx_hstatus; /* health status of the transmit */
ktime_t tx_deadline; /* completion deadline */
u64 tx_cookie; /* completion cookie */
struct lnet_msg *tx_lntmsg[2]; /* lnet msgs to finalize on completion */
@@ -1027,7 +1028,8 @@ struct kib_conn *kiblnd_create_conn(struct kib_peer_ni *peer_ni,
void kiblnd_close_conn_locked(struct kib_conn *conn, int error);
void kiblnd_launch_tx(struct lnet_ni *ni, struct kib_tx *tx, lnet_nid_t nid);
-void kiblnd_txlist_done(struct list_head *txlist, int status);
+void kiblnd_txlist_done(struct list_head *txlist, int status,
+ enum lnet_msg_hstatus hstatus);
void kiblnd_qp_event(struct ib_event *event, void *arg);
void kiblnd_cq_event(struct ib_event *event, void *arg);
diff --git a/net/lnet/klnds/o2iblnd/o2iblnd_cb.c b/net/lnet/klnds/o2iblnd/o2iblnd_cb.c
index 60706b4..007058a 100644
--- a/net/lnet/klnds/o2iblnd/o2iblnd_cb.c
+++ b/net/lnet/klnds/o2iblnd/o2iblnd_cb.c
@@ -89,12 +89,17 @@ static int kiblnd_init_rdma(struct kib_conn *conn, struct kib_tx *tx, int type,
if (!lntmsg[i])
continue;
+ /* propagate health status to LNet for requests */
+ if (i == 0 && lntmsg[i])
+ lntmsg[i]->msg_health_status = tx->tx_hstatus;
+
lnet_finalize(lntmsg[i], rc);
}
}
void
-kiblnd_txlist_done(struct list_head *txlist, int status)
+kiblnd_txlist_done(struct list_head *txlist, int status,
+ enum lnet_msg_hstatus hstatus)
{
struct kib_tx *tx;
@@ -105,6 +110,7 @@ static int kiblnd_init_rdma(struct kib_conn *conn, struct kib_tx *tx, int type,
/* complete now */
tx->tx_waiting = 0;
tx->tx_status = status;
+ tx->tx_hstatus = hstatus;
kiblnd_tx_done(tx);
}
}
@@ -134,6 +140,7 @@ static int kiblnd_init_rdma(struct kib_conn *conn, struct kib_tx *tx, int type,
LASSERT(!tx->tx_nfrags);
tx->tx_gaps = false;
+ tx->tx_hstatus = LNET_MSG_STATUS_OK;
return tx;
}
@@ -265,10 +272,12 @@ static int kiblnd_init_rdma(struct kib_conn *conn, struct kib_tx *tx, int type,
}
if (!tx->tx_status) { /* success so far */
- if (status < 0) /* failed? */
+ if (status < 0) { /* failed? */
tx->tx_status = status;
- else if (txtype == IBLND_MSG_GET_REQ)
+ tx->tx_hstatus = LNET_MSG_STATUS_REMOTE_ERROR;
+ } else if (txtype == IBLND_MSG_GET_REQ) {
lnet_set_reply_msg_len(ni, tx->tx_lntmsg[1], status);
+ }
}
tx->tx_waiting = 0;
@@ -846,6 +855,7 @@ static int kiblnd_map_tx(struct lnet_ni *ni, struct kib_tx *tx,
* posted NOOPs complete
*/
spin_unlock(&conn->ibc_lock);
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
kiblnd_tx_done(tx);
spin_lock(&conn->ibc_lock);
CDEBUG(D_NET, "%s(%d): redundant or enough NOOP\n",
@@ -1045,6 +1055,7 @@ static int kiblnd_map_tx(struct lnet_ni *ni, struct kib_tx *tx,
conn->ibc_noops_posted--;
if (failed) {
+ tx->tx_hstatus = LNET_MSG_STATUS_REMOTE_DROPPED;
tx->tx_waiting = 0; /* don't wait for peer_ni */
tx->tx_status = -EIO;
}
@@ -1393,7 +1404,8 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
CWARN("Abort reconnection of %s: %s\n",
libcfs_nid2str(peer_ni->ibp_nid), reason);
- kiblnd_txlist_done(&txs, -ECONNABORTED);
+ kiblnd_txlist_done(&txs, -ECONNABORTED,
+ LNET_MSG_STATUS_LOCAL_ABORTED);
return false;
}
@@ -1471,6 +1483,7 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
if (tx) {
tx->tx_status = -EHOSTUNREACH;
tx->tx_waiting = 0;
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
kiblnd_tx_done(tx);
}
return;
@@ -1607,6 +1620,7 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
if (rc) {
CERROR("Can't setup GET sink for %s: %d\n",
libcfs_nid2str(target.nid), rc);
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
kiblnd_tx_done(tx);
return -EIO;
}
@@ -1757,6 +1771,7 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
return;
failed_1:
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
kiblnd_tx_done(tx);
failed_0:
lnet_finalize(lntmsg, -EIO);
@@ -1839,6 +1854,7 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
if (rc) {
CERROR("Can't setup PUT sink for %s: %d\n",
libcfs_nid2str(conn->ibc_peer->ibp_nid), rc);
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
kiblnd_tx_done(tx);
/* tell peer_ni it's over */
kiblnd_send_completion(rx->rx_conn, IBLND_MSG_PUT_NAK,
@@ -2050,13 +2066,34 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
if (txs == &conn->ibc_active_txs) {
LASSERT(!tx->tx_queued);
LASSERT(tx->tx_waiting || tx->tx_sending);
+ if (conn->ibc_comms_error == -ETIMEDOUT) {
+ if (tx->tx_waiting && !tx->tx_sending)
+ tx->tx_hstatus =
+ LNET_MSG_STATUS_REMOTE_TIMEOUT;
+ else if (tx->tx_sending)
+ tx->tx_hstatus =
+ LNET_MSG_STATUS_NETWORK_TIMEOUT;
+ }
} else {
LASSERT(tx->tx_queued);
+ if (conn->ibc_comms_error == -ETIMEDOUT)
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_TIMEOUT;
+ else
+ tx->tx_hstatus = LNET_MSG_STATUS_LOCAL_ERROR;
}
tx->tx_status = -ECONNABORTED;
tx->tx_waiting = 0;
+ /* TODO: This makes an assumption that
+ * kiblnd_tx_complete() will be called for each tx. If
+ * that event is dropped we could end up with stale
+ * connections floating around. We'd like to deal with
+ * that in a better way.
+ *
+ * Also that means we can exceed the timeout by many
+ * seconds.
+ */
if (!tx->tx_sending) {
tx->tx_queued = 0;
list_del(&tx->tx_list);
@@ -2066,7 +2103,10 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
spin_unlock(&conn->ibc_lock);
- kiblnd_txlist_done(&zombies, -ECONNABORTED);
+ /* aborting transmits occurs when finalizing the connection.
+ * The connection is finalized on error
+ */
+ kiblnd_txlist_done(&zombies, -ECONNABORTED, -1);
}
static void
@@ -2147,7 +2187,8 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
CNETERR("Deleting messages for %s: connection failed\n",
libcfs_nid2str(peer_ni->ibp_nid));
- kiblnd_txlist_done(&zombies, -EHOSTUNREACH);
+ kiblnd_txlist_done(&zombies, error,
+ LNET_MSG_STATUS_LOCAL_DROPPED);
}
static void
@@ -2223,7 +2264,8 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
kiblnd_close_conn_locked(conn, -ECONNABORTED);
write_unlock_irqrestore(&kiblnd_data.kib_global_lock, flags);
- kiblnd_txlist_done(&txs, -ECONNABORTED);
+ kiblnd_txlist_done(&txs, -ECONNABORTED,
+ LNET_MSG_STATUS_LOCAL_ERROR);
return;
}
@@ -3300,7 +3342,8 @@ static int kiblnd_resolve_addr(struct rdma_cm_id *cmid,
write_unlock_irqrestore(&kiblnd_data.kib_global_lock, flags);
if (!list_empty(&timedout_txs))
- kiblnd_txlist_done(&timedout_txs, -ETIMEDOUT);
+ kiblnd_txlist_done(&timedout_txs, -ETIMEDOUT,
+ LNET_MSG_STATUS_LOCAL_TIMEOUT);
/*
* Handle timeout by closing the whole
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:56 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:56 -0500
Subject: [lustre-devel] [PATCH 008/622] lustre: obdecho: turn on async flag
only for mode 3
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-9-git-send-email-jsimmons@infradead.org>
From: Rahul Deshmukh
There are couple of problems in obdfilter-survey:
- Type of test brw i.e. "g" was not followed with npages,
- Target netdisk was not set properly and
- Turn ON async flag only for mode 3.
This patch fixes the last problem which is kernel side.
WC-bug-id: https://jira.whamcloud.com/browse/LU-5031
Lustre-commit: 9f38647a7b24 ("LU-5031 tests: obdfilter-survey fixes")
Signed-off-by: Rahul Deshmukh
Reviewed-on: http://review.whamcloud.com/10264
Reviewed-by: Cliff White
Reviewed-by: Bob Glossman
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/obdecho/echo_client.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/fs/lustre/obdecho/echo_client.c b/fs/lustre/obdecho/echo_client.c
index ca963bb..3984cb4 100644
--- a/fs/lustre/obdecho/echo_client.c
+++ b/fs/lustre/obdecho/echo_client.c
@@ -1425,7 +1425,7 @@ static int echo_client_brw_ioctl(const struct lu_env *env, int rw,
struct obdo *oa = &data->ioc_obdo1;
struct echo_object *eco;
int rc;
- int async = 1;
+ int async = 0;
long test_mode;
LASSERT(oa->o_valid & OBD_MD_FLGROUP);
@@ -1438,14 +1438,14 @@ static int echo_client_brw_ioctl(const struct lu_env *env, int rw,
/* OFD/obdfilter works only via prep/commit */
test_mode = (long)data->ioc_pbuf1;
- if (test_mode == 1)
- async = 0;
-
if (!ed->ed_next && test_mode != 3) {
test_mode = 3;
data->ioc_plen1 = data->ioc_count;
}
+ if (test_mode == 3)
+ async = 1;
+
/* Truncate batch size to maximum */
if (data->ioc_plen1 > PTLRPC_MAX_BRW_SIZE)
data->ioc_plen1 = PTLRPC_MAX_BRW_SIZE;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:07:58 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:07:58 -0500
Subject: [lustre-devel] [PATCH 010/622] lustre: llite: increase whole-file
readahead to RPC size
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-11-git-send-email-jsimmons@infradead.org>
From: Andreas Dilger
Increase the default whole-file readahead limit to match the current
RPC size. That ensures that files smaller than the RPC size will be
read in a single round-trip instead of sending multiple smaller RPCs.
WC-bug-id: https://jira.whamcloud.com/browse/LU-7990
Lustre-commit: 627d0133d9d7 ("LU-7990 llite: increase whole-file readahead to RPC size")
Signed-off-by: Andreas Dilger
Reviewed-on: https://review.whamcloud.com/26955
Reviewed-by: Patrick Farrell
Reviewed-by: Dmitry Eremin
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/llite/llite_lib.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/fs/lustre/llite/llite_lib.c b/fs/lustre/llite/llite_lib.c
index aaa8ad2..12aafe0 100644
--- a/fs/lustre/llite/llite_lib.c
+++ b/fs/lustre/llite/llite_lib.c
@@ -465,6 +465,12 @@ static int client_common_fill_super(struct super_block *sb, char *md, char *dt)
sbi->ll_dt_exp->exp_connect_data = *data;
+ /* Don't change value if it was specified in the config log */
+ if (sbi->ll_ra_info.ra_max_read_ahead_whole_pages == -1)
+ sbi->ll_ra_info.ra_max_read_ahead_whole_pages =
+ max_t(unsigned long, SBI_DEFAULT_READAHEAD_WHOLE_MAX,
+ (data->ocd_brw_size >> PAGE_SHIFT));
+
err = obd_fid_init(sbi->ll_dt_exp->exp_obd, sbi->ll_dt_exp,
LUSTRE_SEQ_METADATA);
if (err) {
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:04 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:04 -0500
Subject: [lustre-devel] [PATCH 016/622] lustre: ldlm: don't disable softirq
for exp_rpc_lock
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-17-git-send-email-jsimmons@infradead.org>
From: Liang Zhen
it is not necessary to call ldlm_lock_busy() in the context of timer
callback, we can call it in thread context of expired_lock_main.
With this change, we don't need to disable softirq for exp_rpc_lock.
Instead of moving busy locks to the end of the waiting list one
at a time in the context of the timer callback, move any locks
that may be expired onto the expired list. If these locks are
still being used by RPCs being processed, then put them back
onto the end of the waiting list instead of evicting the client.
For the linux client the impact of this change is change of
spin_lock_bh() to spin_lock() for the exp_rpc_lock.
WC-bug-id: https://jira.whamcloud.com/browse/LU-6032
Lustre-commit: 292aa42e0897 ("LU-6032 ldlm: don't disable softirq for exp_rpc_lock")
Signed-off-by: Liang Zhen
Reviewed-on: https://review.whamcloud.com/12957
Reviewed-by: Dmitry Eremin
Reviewed-by: Andreas Dilger
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/ptlrpc/service.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/fs/lustre/ptlrpc/service.c b/fs/lustre/ptlrpc/service.c
index d57df36..3c61e83 100644
--- a/fs/lustre/ptlrpc/service.c
+++ b/fs/lustre/ptlrpc/service.c
@@ -1307,9 +1307,9 @@ static int ptlrpc_server_hpreq_init(struct ptlrpc_service_part *svcpt,
LASSERT(rc <= 1);
}
- spin_lock_bh(&req->rq_export->exp_rpc_lock);
+ spin_lock(&req->rq_export->exp_rpc_lock);
list_add(&req->rq_exp_list, &req->rq_export->exp_hp_rpcs);
- spin_unlock_bh(&req->rq_export->exp_rpc_lock);
+ spin_unlock(&req->rq_export->exp_rpc_lock);
}
ptlrpc_nrs_req_initialize(svcpt, req, rc);
@@ -1327,9 +1327,9 @@ static void ptlrpc_server_hpreq_fini(struct ptlrpc_request *req)
if (req->rq_ops->hpreq_fini)
req->rq_ops->hpreq_fini(req);
- spin_lock_bh(&req->rq_export->exp_rpc_lock);
+ spin_lock(&req->rq_export->exp_rpc_lock);
list_del_init(&req->rq_exp_list);
- spin_unlock_bh(&req->rq_export->exp_rpc_lock);
+ spin_unlock(&req->rq_export->exp_rpc_lock);
}
}
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:07 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:07 -0500
Subject: [lustre-devel] [PATCH 019/622] lustre: hsm: ignore compound_id
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-20-git-send-email-jsimmons@infradead.org>
From: "John L. Hammond"
Ignore request compound ids in the HSM coordinator. Compound ids
prevent batching of CDT to CT requests and degrade HSM
performance. Use CT/archive id compatabiliy when deciding which HSM
actions to put in a request.
WC-bug-id: https://jira.whamcloud.com/browse/LU-10383
Lustre-commit: 9ee81f920bb3 ("LU-10383 hsm: ignore compound_id")
Signed-off-by: John L. Hammond
Reviewed-on: https://review.whamcloud.com/30949
Reviewed-by: Quentin Bouget
Reviewed-by: Faccini Bruno
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
include/uapi/linux/lustre/lustre_idl.h | 2 +-
include/uapi/linux/lustre/lustre_user.h | 2 +-
2 files changed, 2 insertions(+), 2 deletions(-)
diff --git a/include/uapi/linux/lustre/lustre_idl.h b/include/uapi/linux/lustre/lustre_idl.h
index 4e1605a2..307feb3 100644
--- a/include/uapi/linux/lustre/lustre_idl.h
+++ b/include/uapi/linux/lustre/lustre_idl.h
@@ -2508,7 +2508,7 @@ struct llog_agent_req_rec {
*/
__u32 arr_archive_id; /**< backend archive number */
__u64 arr_flags; /**< req flags */
- __u64 arr_compound_id;/**< compound cookie */
+ __u64 arr_compound_id;/**< compound cookie, ignored */
__u64 arr_req_create; /**< req. creation time */
__u64 arr_req_change; /**< req. status change time */
struct hsm_action_item arr_hai; /**< req. to the agent */
diff --git a/include/uapi/linux/lustre/lustre_user.h b/include/uapi/linux/lustre/lustre_user.h
index 27501a2..5405e1b 100644
--- a/include/uapi/linux/lustre/lustre_user.h
+++ b/include/uapi/linux/lustre/lustre_user.h
@@ -1729,7 +1729,7 @@ static inline char *hai_dump_data_field(struct hsm_action_item *hai,
struct hsm_action_list {
__u32 hal_version;
__u32 hal_count; /* number of hai's to follow */
- __u64 hal_compound_id; /* returned by coordinator */
+ __u64 hal_compound_id; /* returned by coordinator, ignored */
__u64 hal_flags;
__u32 hal_archive_id; /* which archive backend */
__u32 padding1;
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:09 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:09 -0500
Subject: [lustre-devel] [PATCH 021/622] lustre: ptlrpc:
ptlrpc_register_bulk() LBUG on ENOMEM
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-22-git-send-email-jsimmons@infradead.org>
From: Andriy Skulysh
Assertion fails on !desc->bd_registered during
retry after ENOMEM.
Drop bd_registered flag and exit via cleanup_bulk
to ensure that bulk is fully unregistered.
Cray-bug-id: MRP-4733
WC-bug-id: https://jira.whamcloud.com/browse/LU-10643
Lustre-commit: 4a81be263079 ("LU-10643 ptlrpc: ptlrpc_register_bulk() LBUG on ENOMEM")
Signed-off-by: Andriy Skulysh
Reviewed-on: https://review.whamcloud.com/31228
Reviewed-by: Alexandr Boyko
Reviewed-by: Andrew Perepechko
Reviewed-by: Oleg Drokin
Signed-off-by: James Simmons
---
fs/lustre/include/obd_support.h | 1 +
fs/lustre/ptlrpc/niobuf.c | 12 +++++++++---
2 files changed, 10 insertions(+), 3 deletions(-)
diff --git a/fs/lustre/include/obd_support.h b/fs/lustre/include/obd_support.h
index 653a456..67500b5 100644
--- a/fs/lustre/include/obd_support.h
+++ b/fs/lustre/include/obd_support.h
@@ -349,6 +349,7 @@
#define OBD_FAIL_PTLRPC_DROP_BULK 0x51a
#define OBD_FAIL_PTLRPC_LONG_REQ_UNLINK 0x51b
#define OBD_FAIL_PTLRPC_LONG_BOTH_UNLINK 0x51c
+#define OBD_FAIL_PTLRPC_BULK_ATTACH 0x521
#define OBD_FAIL_OBD_PING_NET 0x600
#define OBD_FAIL_OBD_LOG_CANCEL_NET 0x601
diff --git a/fs/lustre/ptlrpc/niobuf.c b/fs/lustre/ptlrpc/niobuf.c
index 02ed373..2e866fe 100644
--- a/fs/lustre/ptlrpc/niobuf.c
+++ b/fs/lustre/ptlrpc/niobuf.c
@@ -179,8 +179,13 @@ static int ptlrpc_register_bulk(struct ptlrpc_request *req)
LNET_MD_OP_GET : LNET_MD_OP_PUT);
ptlrpc_fill_bulk_md(&md, desc, posted_md);
- rc = LNetMEAttach(desc->bd_portal, peer, mbits, 0,
- LNET_UNLINK, LNET_INS_AFTER, &me_h);
+ if (posted_md > 0 && posted_md + 1 == total_md &&
+ OBD_FAIL_CHECK(OBD_FAIL_PTLRPC_BULK_ATTACH)) {
+ rc = -ENOMEM;
+ } else {
+ rc = LNetMEAttach(desc->bd_portal, peer, mbits, 0,
+ LNET_UNLINK, LNET_INS_AFTER, &me_h);
+ }
if (rc != 0) {
CERROR("%s: LNetMEAttach failed x%llu/%d: rc = %d\n",
desc->bd_import->imp_obd->obd_name, mbits,
@@ -209,6 +214,7 @@ static int ptlrpc_register_bulk(struct ptlrpc_request *req)
LASSERT(desc->bd_md_count >= 0);
mdunlink_iterate_helper(desc->bd_mds, desc->bd_md_max_brw);
req->rq_status = -ENOMEM;
+ desc->bd_registered = 0;
return -ENOMEM;
}
@@ -585,7 +591,7 @@ int ptl_send_rpc(struct ptlrpc_request *request, int noreply)
if (request->rq_bulk) {
rc = ptlrpc_register_bulk(request);
if (rc != 0)
- goto out;
+ goto cleanup_bulk;
/*
* All the mds in the request will have the same cpt
* encoded in the cookie. So we can just get the first
--
1.8.3.1
From jsimmons at infradead.org Thu Feb 27 21:08:11 2020
From: jsimmons at infradead.org (James Simmons)
Date: Thu, 27 Feb 2020 16:08:11 -0500
Subject: [lustre-devel] [PATCH 023/622] lustre: osc: Do not request more
than 2GiB grant
In-Reply-To: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
References: <1582838290-17243-1-git-send-email-jsimmons@infradead.org>
Message-ID: <1582838290-17243-24-git-send-email-jsimmons@infradead.org>
From: Patrick Farrell
The server enforces a grant limit of 2 GiB, which the
client must honor. The existing client code combined with
16 MiB RPCs make it possible for the client to ask for
more than this limit.
Make this limit explicit, and also fix an overflow bug in
o_undirty calculation in osc_announce_cached. (o_undirty
is a 32 bit value and 16 MiB*256 rpcs_in_flight = 4 GiB.
4 GiB + extra grant components overflows o_undirty.)
Cray-bug-id: LUS-5750
WC-bug-id: https://jira.whamcloud.com/browse/LU-10776
Lustre-commit: c0246d887809 ("LU-10776 osc: Do not request more than 2GiB grant")
Signed-off-by: Patrick Farrell
Reviewed-on: https://review.whamcloud.com/31533
Reviewed-by: Nathaniel Clark
Reviewed-by: Bobi Jam
Reviewed-by: Andrew Perepechko